Appendix B

The Fast-Start Catalog: Twelve Pre-Scoped Briefs

Twelve ready-to-build projects, each with a real problem, a named user, drafted starter requirements, the genuinely hard part, and a stack already sized to 160 hours

Appendix B — The Fast-Start Catalog: Twelve Pre-Scoped Briefs

“Whatever your hand finds to do, do it with your might.” — Ecclesiastes 9:10 (ESV)

This appendix is the reason this course is eight weeks long instead of sixteen.

An accelerated capstone is not lost in the build. It is lost in the four weeks a student spends deciding what to build — reading, browsing, half-starting, waiting for an idea worthy of the effort. That deliberation is a real skill, and the sixteen-week edition spends two full weeks teaching it. You do not have two weeks. So this catalog removes the blank page: twelve projects that a senior can actually finish in about 160 hours, each one already carrying the work that normally eats Weeks 1 through 4.

Every brief ships the same eight things: the problem and a specific person who has it; a minimum viable scope of three to four Must features; eight to ten drafted functional requirements with acceptance criteria; four to six non-functional requirements with measurable targets; the genuinely hard part that will eat your schedule; a stack that fits, with the reasoning; what to cut first if Week 4 says you are behind; and one way to go further if you are ahead.


B.0 — How to choose in ninety minutes

Not ninety minutes of browsing. Ninety minutes, timed, ending in a decision you write down.

MinutesDo this
0–20Read the One line and The person who has this problem for all twelve. Nothing else. Strike any brief whose user you cannot picture caring.
20–45You should have three or four left. Read their genuinely hard part section — only that section. This is the single most predictive paragraph in the brief, because the hard part is what will still be unfinished in Week 7. Strike anything whose hard part you cannot describe to another person.
45–70Two left. Read both in full, including the stack and the cut list. Ask the only question that matters at this stage: which of these can I get a walking skeleton of by the end of Week 3?
70–90Decide. Write the decision and the reason into docs/scoping-decision.md. Name the brief, name what you are changing about it, and name the one you rejected and why.

Three rules for the ninety minutes:

  1. Pick the one you can finish, not the one you would brag about. The impressive project you abandon in Week 6 scores worse than the modest one you hand off cleanly in Week 8. The rubric grades a finished, documented, reproducible system — it has no line for ambition.
  2. The hard part is the project. Every brief names one. If two briefs appeal equally, take the one whose hard part teaches you something you actually want to know.
  3. Do not blend two briefs. It is the most common Week-1 mistake and it silently doubles the scope, because you inherit both hard parts and neither cut list.

Coach’s Note — If you already have a project — one from your job, your research, a client, a genuine itch — take it. You are not required to use this catalog. Use the shape instead: open the brief closest to your domain and write your own project in that format, section for section. The twenty hours a brief saves you drop to about eight, but the format is doing most of the work.


B.1 — Adopting is a decision, not a default

Adopting a brief is not copying it. The moment you choose one, you inherit a set of statements written for a hypothetical student — and an inherited requirement that quietly describes somebody else’s project is the signature failure of this edition. Chapter 2 trains the repair. The short version, applied to every requirement you inherit:

  • Does my actor exist? A brief says “household member.” Your project may have club members, staff, or one person. If the actor changed, the requirement changed.
  • Does my scope still contain it? If you cut a feature in Week 1, every requirement whose condition depends on it is withdrawn — including the one that only mentions it in an acceptance criterion.
  • Does the acceptance data exist in my project? “Given two members with notifications enabled” is not a testable criterion if you have no notifications.

Four verdicts, applied to every inherited requirement: keep, adapt, split, or withdraw. Record the verdict. The adaptation record is a graded part of Milestone 2, and it is the thing that proves the specification is yours rather than the catalog’s.


B.2 — The twelve briefs

Read the table, then the brief.

#BriefShapeThe hard part, in a phrase
1PantryPilotWeb app, auth, third-party lookupExpiry rules that survive real messy input
2TraceLensCommand line, data processingParsing formats you do not control
3SlotCheckWeb app, schedulingConflict detection across overlapping rules
4TrendDeckData ingest and dashboardMaking a public dataset trustworthy
5GroundedApplied AI, retrievalRefusing to answer when support is absent
6Repo HealthDeveloper tooling, CLITurning judgment into defensible evidence
7Own AuditSecurity, defensive onlyRedaction, and proving you had authority
8Access CheckAccessibility toolingThe checks automation cannot make
9Shift DeckCivic/nonprofit web appA real client with real constraints
10IntervalEducation technologyScheduling that survives an honest user
11Round TripOffline-first web appSync conflicts with no server to arbitrate
12Restore FirstSystems toolingProving a restore actually restored

Two of these — PantryPilot and TraceLens — are the running examples used throughout this book and its sixteen-week edition, so every worked example you meet elsewhere is an example of a project in this catalog.


Brief 1 — PantryPilot

One line: A small web application that lets a roommate group track what food the household has, see what is about to spoil, and decide what to cook from it tonight.

The person who has this problem: Dana, 22, the organizer in a four-person apartment. She buys most of the groceries and keeps a whiteboard list that goes stale in about four days. Her housemate Marcus, 21, will never enter data and will only open something that tells him what to eat. Dana is the running example’s user in the sixteen-week edition; she is not yours. Your first job in Week 1 is to find a real household — yours counts if you actually live in it — and replace Dana with a person you can name, on a date you can write down.

Why it fits 160 hours: A vertical slice costs seven to nine hours in a stack you already know. The Must list below prices out bottom-up at about 27 hours of feature construction — keep the pantry current is three small slices, not one — which is what the 40-hour build-and-verify window in Weeks 5–6 actually buys once you reserve roughly 12 of those hours for tests, integration, and the defect log. Add about 8 hours of walking skeleton and continuous integration out of Week 3’s twenty, and about 6 hours of deployment out of Week 7’s twenty. That is roughly 41 construction-shaped hours out of 160. The other 119 are the charter, the requirements, the design, the review, the documentation, the handoff, and the delivery — and none of that gets cut.

The problem

Four people share a kitchen and nobody shares a memory of it. Somebody buys sour cream on Tuesday because nobody could remember whether there was sour cream. Two weeks later somebody else throws out two containers of sour cream, one of which was never opened. The tortillas at the back of the second shelf go green. In a house like Dana’s this happens quietly and continuously, and the money it costs is real but invisible, because nobody ever sees the total — they see one container at a time, in the trash, on a Thursday.

Dana already tried to solve it. There is a whiteboard by the door with a list on it, and the list is right for about four days after she writes it. It goes wrong in a specific way that matters: people cross things off when they finish them, but they never add things, and nobody writes down dates at all. So the whiteboard tells you approximately what was bought last week, which is not the question anyone is asking. The question is what is about to go bad, and the whiteboard has never been able to answer it.

Marcus is the harder half of the problem, and he is the reason this project is not just a shared to-do list. He will not type. He will not maintain anything. If the system depends on every housemate entering every item, it will die in nine days the way the whiteboard died. So the design has to work when only one person maintains the data and the rest of the house consumes it — which is a real constraint with real consequences, and it is the sort of thing you would never discover from a feature list.

Minimum viable scope (the Must list)

  • Enter the household. Any housemate joins the shared pantry by typing a six-character code. No accounts, no passwords, no email — a session on that device, scoped to that one household. (~9 h)
  • Keep the pantry current. Add an item with a name, quantity, unit, and expiry date; see the whole household’s list; mark an item consumed so it leaves the list. (~12 h)
  • See what is about to spoil. A view of every item expiring inside a configurable window, soonest first, with a real empty state. (~6 h)
  • Run somewhere that is not your laptop. Deployed, seeded with realistic data, and demonstrable to a stranger in ten minutes. (~6 h, and they come out of Week 7, not out of the build window)

That is three feature slices and a deployment, and it is the whole project. The sixteen-week edition of this book carries five Musts — it adds barcode entry at fourteen hours — because it has 240 hours to spend. You have 160. Barcode lookup and the AI recipe suggestion are both in this brief, and you can afford at most one of them, after Week 4 tells you that you are ahead. Adopting both is the single most common way an accelerated student ships a broken demo. Per-person accounts, the daily expiring-soon notification, waste tracking, a shopping list, and photographs of items are not in the release.

Starter functional requirements

Ten drafted requirements. Every one of them is mine, which means every one of them is missing the thing the house style requires: a Source: line naming a human, an observation, or a decision you recorded. Add one to each requirement you keep, or delete the requirement.

FR-ACC-01 — Join a household with a code

Priority: Must Requirement: A person shall be able to enter a household by supplying its six-character join code, receiving a session scoped to that one household and lasting fourteen days. Rationale: The household is the unit of sharing. Replacing accounts, passwords, and invitations with one shared code is a scoping decision, not a shortcut: anyone holding the code has full access and there is no per-person history, and that cost is accepted in writing. Acceptance criteria:

  • Given a valid, current join code, when a person submits it, then they receive a session bound to exactly one household and see that household’s pantry.
  • Given an invalid code, when a person submits it, then the system returns a single generic error that does not reveal whether the code exists, and the attempt counts against the rate limit in NFR-SEC-02.
  • Given a session older than fourteen days, when the person opens the application, then they are returned to the join screen with their pantry data intact.

FR-INV-01 — Add a pantry item

Priority: Must Requirement: A household member with an active session shall be able to add an item to the household pantry by supplying a name, a quantity with a unit, and an expiry date. Rationale: Nothing else in the system works until inventory exists. Dana’s whiteboard is the behavior being replaced. Acceptance criteria:

  • Given a member on the pantry screen, when they submit a name, quantity, unit, and expiry date, then the item appears in the household pantry list within one page refresh and is visible to every member of that household.
  • Given a submission with an expiry date earlier than today, when the member submits, then the item is saved and displayed in the expired group rather than rejected.
  • Given a submission missing the item name, when the member submits, then the system rejects the submission and states which field is missing.

FR-INV-03 — List the household pantry

Priority: Must Requirement: A household member with an active session shall be able to view every active item in the household pantry, showing name, quantity with unit, and expiry date. Rationale: The list is the shared memory the whiteboard failed to be. It is also the screen Marcus will look at, so it has to be readable without instruction. Acceptance criteria:

  • Given a household with items entered by two different members, when any member opens the pantry list, then every active item from both members appears.
  • Given a household with no items, when a member opens the pantry list, then the system shows an explicit empty state naming the next action, not a blank screen.

FR-INV-04 — Mark an item consumed

Priority: Must Requirement: A household member with an active session shall be able to mark a pantry item as consumed, which removes it from the active pantry list. Rationale: An inventory that only grows is worse than no inventory. Crossed-out lines on the whiteboard were never erased, and within a week nobody trusted the board. Acceptance criteria:

  • Given an item in the active pantry list, when a member marks it consumed, then it no longer appears in the active list and is not counted by the expiring-soon view.
  • Given an item marked consumed in error, when the member selects undo within the same session, then the item returns to the active list with its original expiry date.

FR-EXP-02 — Expiring-soon view

Priority: Must Requirement: A household member with an active session shall be able to view every pantry item whose expiry date falls within a household-configured window, ordered soonest first, with already-expired items marked. Rationale: This is the whole reason Dana would open the application. If this screen is wrong, nothing else in the project matters. Acceptance criteria:

  • Given a seven-day window and a pantry containing items expiring in 2, 6, and 20 days, when a member opens the expiring-soon view, then exactly the 2-day and 6-day items are listed, in that order.
  • Given a pantry with no items expiring inside the window, when a member opens the view, then the system displays an explicit empty state rather than a blank screen.
  • Given an item that expired three days ago, when a member opens the view, then the item appears and is visibly marked as already expired.

FR-INV-02 — Edit an item

Priority: Should Requirement: A household member with an active session shall be able to change a pantry item’s quantity, unit, or expiry date. Rationale: People buy two and use one. Without an edit path the only way to correct anything is delete-and-retype, which is exactly the friction that killed the whiteboard. Acceptance criteria:

  • Given an item in the active list, when a member changes its quantity and saves, then the new quantity is visible to every member of the household within one page refresh.
  • Given a member changes an expiry date so the item now falls inside the expiring-soon window, when they save, then the item appears in the expiring-soon view without any further action.

FR-SCAN-01 — Add an item by barcode

Priority: Should Requirement: A household member with an active session shall be able to add a pantry item by submitting a product barcode, which the system uses to pre-fill the item name and unit. Rationale: Typing every item is the reason the whiteboard died, and Marcus will scan before he will type. This is also the most expensive Should in the brief — budget fourteen hours, not four. Acceptance criteria:

  • Given a barcode the product-lookup service recognizes, when the member submits it, then the add-item form opens with the product name and unit pre-filled and the quantity and expiry fields empty.
  • Given a barcode the service does not recognize, when the member submits it, then the system opens the manual add-item form with the barcode retained and states that no product was found.

FR-SCAN-02 — Behavior when the lookup service is unavailable

Priority: Must if FR-SCAN-01 is adopted; otherwise Won’t (this release) Requirement: When the product-lookup service does not respond within five seconds, the system shall present the manual add-item form with any data the member has already entered preserved. Rationale: The one third-party dependency in this project will be down at some point, and the most likely moment is your Week-8 demo. A feature with no defined behavior under failure is a feature that fails in front of an audience. Acceptance criteria:

  • Given the lookup service is unreachable, when a member submits a barcode, then within six seconds the manual form appears, the barcode is still in the field, and a message states that lookup is unavailable.
  • Given the lookup service is unreachable, when a member completes the manual form, then the item saves normally and no error is surfaced afterwards.

FR-REC-01 — Cooking suggestion from what is on hand

Priority: Could Requirement: Given at least five items in the household pantry, the system shall return between one and three cooking suggestions within eight seconds, each listing the pantry items it uses and any ingredients not on hand. Rationale: This is Marcus’s only reason to open the application, and it is the feature every reviewer asks about. It is also eighteen hours, an evaluation set, a fallback, and a spend cap, which is why it is a Could in a 160-hour budget and not a Should. Acceptance criteria:

  • Given a pantry of at least five items, when a member requests a suggestion, then between one and three suggestions are returned within eight seconds, each naming at least three items that are actually in the pantry, each clearly labeled as generated.
  • Given the twenty-case evaluation set committed to the repository, when the suggester is run against it, then at least sixteen cases satisfy the envelope above, with zero cases presenting a missing ingredient as on hand.
  • Given the model is unavailable or has not responded within eight seconds, when a member requests a suggestion, then the system shows the three items closest to expiry, states that no suggestion is available, and says why.

FR-EXP-05 — Daily expiring-soon notification

Priority: Won’t (this release) Requirement: The system shall send one notification per household per day, at a household-configured time, listing every item expiring within the next three days. Rationale: Recorded so the decision is visible rather than forgotten. PantryPilot has no accounts and no email addresses, so there is no delivery address to send to; adding one means adding identity, which is the cost the join code was chosen to avoid. The sixteen-week edition can afford this as a Should. You cannot. Revisit if a v1.0 ships with time left in Week 8. A Won’t requirement needs no acceptance criteria — there is nothing to accept.

Count the priorities before you move on: five Musts, two Shoulds, one Could, one Won’t, one conditional. If your edited version comes back with nine Musts, you have not prioritized — you have relabeled “everything” in project-management vocabulary, and Week 6 will find out for you.

Starter non-functional requirements

NFR-PERF-01 — Pantry list render time

Priority: Must Requirement: The pantry list view renders at a 95th percentile under 1.5 seconds with 200 seeded items, on a throttled Fast 3G profile with a cold cache. Measured by: twenty loads in browser developer tools with throttling applied; the p95 recorded in the measurements log in Week 6 and again in Week 8.

NFR-SEC-01 — No secrets in the repository

Priority: Must Requirement: No credential, API key, or token appears in the repository at any commit in its history. Measured by: a secret-scanning step in continuous integration over full history, returning zero findings on the release commit. Configuration comes from environment variables; a placeholder example file is committed and the real one is ignored.

NFR-SEC-02 — Join codes cannot be guessed

Priority: Must Requirement: A join code is drawn from a cryptographically secure random source, is stored only as a salted hash produced by a current, well-reviewed password-hashing function, and never appears in a log; code entry is limited to 10 attempts per hour per client. Measured by: a unit test asserting the stored value is not the plaintext code; a test asserting the eleventh attempt within an hour is refused; a grep of the captured log fixture for a known code, returning nothing.

NFR-REL-01 — The deployed system stays up

Priority: Must Requirement: The deployed application answers a health check successfully on at least 13 of 14 consecutive daily checks across the final two weeks. Measured by: a scheduled check that appends pass or fail with a timestamp to the measurements log. Note that the sixteen-week edition asks for 29 of 30 across four weeks; you deploy in Week 7, so you have two weeks of evidence, not four. Adapting a threshold to a real measurement window is exactly the skill this catalog is testing.

NFR-ACC-01 — Keyboard and contrast

Priority: Must Requirement: Every interactive control is reachable and operable by keyboard alone in a logical order with a visible focus indicator; body text meets a contrast ratio of at least 4.5:1 against its background; no information is conveyed by color alone. Measured by: unplug the mouse and complete the three core tasks, noting every place you got stuck; run a contrast checker over every text color pair; set the display to grayscale and repeat the three tasks.

NFR-PRIV-01 — What leaves the household

Priority: Should Requirement: No data other than the barcode string leaves the system to the product-lookup service, and — if FR-REC-01 is adopted — no data other than item names and expiry dates leaves the system to the model provider. No member name, join code, or device identifier is ever included. Measured by: a unit test over the outbound request builder asserting the exact field set; a manual capture of one real request, redacted and pasted into the requirements document with the date.

The genuinely hard part

It is not the forms. It is the expiry model and everything that hangs off a date. Half the food in a real pantry has no printed date, a third has a date that means “best by” rather than “unsafe after,” rice does not expire in any useful sense, and an opened jar expires on a different schedule than a sealed one. The moment you write FR-EXP-02 you have committed to answering: what does the system do with an item that has no date? Does “expiring soon” mean the same thing for milk and for flour? Whose timezone decides when today ends, when one housemate is a night-shift nurse? Every one of those is a decision you will make either in Week 2, in writing, in ten minutes, or in Week 6, in code, under pressure, wrongly.

The second eater is the third-party integration, and it eats in a way students do not expect. Barcode lookup does not fail cleanly — it degrades. Some products are missing entirely. Store-brand items are missing more often than national brands, which means the failure is concentrated exactly where a student household shops. The response almost never contains an expiry date, so the field you most wanted still has to be typed. And every one of those partial results has to reach the user as something better than a spinner. That is why FR-SCAN-02 is a Must and FR-SCAN-01 is only a Should: the fallback is worth more than the feature.

A stack that fits

Node with PostgreSQL, reached through a single repository module, with schema changes shipped as numbered migrations applied by script/setup. That is the sixteen-week edition’s ADR-0001, and the reasoning is worth repeating because it is not the reasoning most students expect: a decision matrix in that book actually scores SQLite higher, and the student chooses PostgreSQL anyway and writes down why — one managed instance, real concurrent writes from four housemates, and a documented dump-and-restore path for the handoff. A matrix is decision support. It is not the decision.

Watch your novelty load. If you have used Node and have written SQL in coursework but have never run PostgreSQL yourself, that is one genuinely new thing and you are fine. Add a front-end framework you have never shipped and you are at two, which is the ceiling this course sets. Add a third — an ORM you have not used, a hosting platform you have never deployed to, a container workflow you are learning as you go — and the arithmetic stops working, because you cannot estimate an error you have not seen yet.

Say it plainly: this is a recommendation, not a verdict. If you are faster in Python, in C#, in Go, or in a Java stack you have shipped before, use it and beat me by twenty hours. What the course requires is not this stack; it is an architecture decision record in Week 3 that names the alternatives you actually considered, the criteria you weighed, and the cost you accepted. Adopting ADR-0001 unchanged is a legitimate decision — but you still have to write the record, in your own words, with your own reasons.

What to cut first if Week 4 says you are behind

  1. Cut FR-REC-01, the cooking suggestion, entirely. Saves 18 hours. Not “simplify it” — cut it, mark it Won’t with a dated change-log row, and delete the prompt experiments from the branch. This is the first cut and it is not close.
  2. Cut FR-SCAN-01 and withdraw FR-SCAN-02 with it. Saves 14 hours. Manual entry was always the primary path; the barcode was a convenience. Withdrawing both together is the correct move, because a fallback with nothing to fall back from is dead code.
  3. Cut FR-INV-02, editing, down to delete-and-re-add. Saves 4 hours. It is worse for the user and honest to say so in the specification.
  4. Fix the expiry window at seven days instead of configurable. Saves 3 hours of settings screen, persistence, and validation. Record it as a Won’t row with a revisit condition, not as a silent deletion.

If you are ahead

Add household waste tracking: when an item is marked discarded rather than consumed, record it, and show one screen with the count and estimated value of what the household threw away in the last thirty days. It closes the loop back to the actual problem — invisible loss — and it is a genuine analytics slice with aggregation queries, a chart, and seed data. Budget 11 hours, and only start it if Week 6 ends with your Musts done, tested, and deployed. Not before.


Brief 2 — TraceLens

One line: A command-line tool that reads a server log file, flags the intervals where something anomalous happened, and writes a report you can read at two in the morning.

The person who has this problem: Priya, an on-call site reliability engineer, at 2 a.m., on a laptop, angry. That is the whole persona list and it is deliberately one line long, because every design decision in this project falls out of it: she is tired, she is in a hurry, she is in a terminal, and she does not want to learn anything. Your real Priya is whoever near you actually keeps logs — a campus systems administrator, the person who runs the department’s web server, a friend on an operations team, or you, on a machine you own. Find them in Week 1 and ask for one thing: a real log file, or permission to point the tool at one.

Why it fits 160 hours: A vertical slice costs seven to nine hours in a stack you already know; these four Musts price out bottom-up at 9, 9, 6, and 5 — about 29 hours of feature construction inside the 40-hour build-and-verify window, leaving roughly 11 for tests and the defect log. The walking skeleton is cheaper here than in a web project — no browser, no server, no authentication — call it 6 to 8 hours of Week 3’s twenty, and packaging plus install verification about 6 hours of Week 7’s twenty. About 43 construction-shaped hours out of 160.

The problem

Something went wrong on the server between one and two in the morning, and the evidence is a log file. Not a dashboard, not a metrics system with a query language — a file, because the thing that broke was on a small deployment nobody funded observability for, which describes most of the machines in the world. Priya opens it and there are 4.2 million lines in it. She knows how to use grep. She has been using grep for eleven minutes and what she has learned is that there are a lot of lines containing 500.

The question she is actually asking is not “which lines say 500.” It is “when did this stop being normal?” That is a different question, and no amount of grep answers it, because normal is not a constant — it is whatever the last hour looked like. Traffic at 2 a.m. is not traffic at 2 p.m. A hundred errors an hour might be Tuesday’s background noise or Wednesday’s outage, and the only way to know is to compare the interval to the intervals around it. That comparison is arithmetic nobody wants to do by hand at 2 a.m., and it is the entire product.

There is a second half to the problem, and it is the half that decides whether the tool is used twice. Real log files are dirty. Lines are truncated mid-write. Two formats are interleaved because someone changed the logger in March and never backfilled. There is a line with a raw newline inside a quoted field. A tool that halts on the first line it does not understand is a tool Priya deletes at 2:04 a.m., because the one thing she cannot afford is a program that fails on her data and tells her nothing about why.

Minimum viable scope (the Must list)

  • Parse two named formats without dying. tracelens scan <file> reads Common Log Format and newline-delimited JSON, and writes every line it cannot parse to a rejects file with its original line number, continuing to the end. (~9 h)
  • Detect one anomaly, defensibly. Bucket records into one-minute intervals and flag every interval whose error count is far enough above the trailing baseline to be worth waking up for, by a rule you can state in one sentence and defend in Week 8. (~9 h)
  • Report in two shapes. Human-readable text on stdout, machine-readable output behind a flag, and documented exit codes so the tool composes with automation that already exists. (~6 h)
  • Survive a file bigger than memory. Stream it. Never read the whole log into a list. (~5 h)

Four slices, no user interface, no accounts, no HTTP, no deployment target beyond “install it and run it.” That is the point of this brief: it is the contrasting shape. Almost all of its difficulty lives in the data, and almost none of it lives in the presentation — which makes it an excellent choice if you are strong at algorithms and a poor choice if what you want to show a hiring panel is a screen. A second detection rule, live tailing of a growing file, alerting, a web interface, and compressed or rotated input are not in the release.

Starter functional requirements

Nine drafted requirements. As with every brief in this catalog, each one needs a Source: line from you before it earns a place in docs/requirements.md.

FR-PARSE-01 — Parse the named formats and reject the rest

Priority: Must Requirement: The tool shall parse log files in Common Log Format and in newline-delimited JSON, writing any line it cannot parse to a rejects file with its original line number, and continuing without terminating. Rationale: Naming the formats converts an infinite obligation into a finite one. The rejects behavior is the part students forget and the part that decides whether the tool is usable at 2 a.m. Acceptance criteria:

  • Given a 10,000-line file with 12 malformed lines, when the tool runs, then 9,988 records are parsed and the rejects file contains 12 entries with their original line numbers.
  • Given a file whose every line is malformed, when the tool runs, then the tool exits with the documented “no records parsed” code, writes all lines to rejects, and prints a message naming the two formats it does understand.

FR-PARSE-02 — Detect the format rather than demand it

Priority: Should Requirement: The tool shall determine each line’s format from its content, so that a file containing both supported formats parses completely without the operator declaring a format. Rationale: The interleaved-formats file is the real case, not the exotic one. It happens every time somebody changes a logger without backfilling. Acceptance criteria:

  • Given a file whose first 500 lines are Common Log Format and whose remaining 500 are newline-delimited JSON, when the tool runs, then 1,000 records are parsed and the rejects file is empty.
  • Given a --format flag naming one format explicitly, when the tool runs, then detection is bypassed and lines in the other format are rejected with their line numbers.

FR-DETECT-01 — Flag anomalous error intervals

Priority: Must Requirement: The tool shall report every one-minute interval in which the count of HTTP 5xx responses exceeds three standard deviations above the mean 5xx count for the preceding sixty minutes. Rationale: “Anomaly” has to be given an operational definition before anything can be built or tested. This rule is a genuine engineering decision, not a default — and because it is a decision, you owe a rationale and an architecture decision record for it. Acceptance criteria:

  • Given a synthetic log with a steady baseline of two 5xx per minute for ninety minutes and a single minute containing forty, when the tool runs, then exactly that one interval is reported, with its timestamp, its count, and the baseline it was compared against.
  • Given a log whose first sixty minutes contain no baseline, when the tool runs, then intervals inside the warm-up window are reported as “insufficient baseline” rather than silently skipped or silently flagged.

FR-DETECT-02 — State the threshold in the output

Priority: Must Requirement: For every reported interval, the tool shall state the observed count, the baseline mean, the baseline standard deviation, and the threshold that was crossed. Rationale: Priya will not act on a tool that says “anomaly” and nothing else, and she should not. A detector that cannot show its work is a detector nobody trusts twice. Acceptance criteria:

  • Given any reported interval, when the operator reads the report, then the four numbers appear on the same record as the timestamp.
  • Given an interval reported as “insufficient baseline”, when the operator reads the report, then the record states how many minutes of history were available.

FR-REPORT-01 — Human-readable report

Priority: Must Requirement: The tool shall write a report to standard output listing every flagged interval in chronological order, preceded by a summary line stating records parsed, records rejected, intervals examined, and intervals flagged. Rationale: The summary line is the first thing Priya reads and often the only thing. If parsing quietly rejected 60% of the file, that fact has to be visible before any conclusion is drawn from the report. Acceptance criteria:

  • Given a run over a file with rejected lines, when the report is written, then the summary line names the reject count and the path to the rejects file.
  • Given a run that flags nothing, when the report is written, then the tool prints the summary and an explicit “no intervals flagged” line, not empty output.

FR-REPORT-02 — Machine-readable output

Priority: Should Requirement: The tool shall write the same report as newline-delimited JSON when --format json is supplied, with one object per flagged interval and one summary object. Rationale: The tool has to compose with the automation Priya already has. A tool that only prints prose is a tool that gets re-parsed with a fragile regular expression by the next person. Acceptance criteria:

  • Given --format json, when the tool runs over the fixture, then every output line parses as valid JSON and the field set matches the schema committed to the repository.
  • Given --format json, when the tool runs, then no human-formatted prose appears on standard output; progress and warnings go to standard error.

FR-CLI-01 — Documented exit codes

Priority: Must Requirement: The tool shall exit 0 when it completes with no intervals flagged, 1 when it completes with one or more intervals flagged, and 2 when it cannot complete, and shall document these three codes in the README. Rationale: Exit codes are the command-line equivalent of an error envelope: decided once, in the specification, so that every path in the program agrees. They are also what lets somebody put the tool in a cron job without reading its source. Acceptance criteria:

  • Given a fixture with a known anomaly, when the tool runs, then the process exits 1.
  • Given a path that does not exist, when the tool runs, then the process exits 2 and standard error names the path.

FR-CLI-02 — Time-range filter

Priority: Should Requirement: The tool shall restrict analysis to records whose timestamps fall inside the range given by --since and --until, when either is supplied. Rationale: Priya knows roughly when it happened. Making her analyze nine days to look at forty minutes wastes the one resource she has none of. Acceptance criteria:

  • Given --since and --until bounding a window containing a known anomaly, when the tool runs, then only intervals inside that window are examined and the summary states the window.
  • Given a --since later than --until, when the tool runs, then the tool exits 2 with a message naming both values.

FR-DETECT-03 — A second detection rule

Priority: Won’t (this release) Requirement: The tool shall detect sustained latency degradation in addition to error bursts. Rationale: Recorded so the decision is visible. A second detector is not a second copy of the first one — it needs its own baseline model, its own threshold defense, its own fixtures, and its own section of the report, and that is fifteen hours the budget does not have. Revisit only if the first detector is finished, tested, and documented by the end of Week 6.

Starter non-functional requirements

NFR-PERF-01 — Throughput on a large file

Priority: Should Requirement: The tool shall process a 1 GB log file in under 120 seconds on the reference machine named in the technical specification. Measured by: three timed runs over a generated 1 GB fixture with the shell’s timing built-in; the median recorded in the measurements log. Name the machine — processor, memory, and disk type — or the number means nothing.

NFR-PERF-02 — Bounded memory

Priority: Must Requirement: Peak resident memory shall stay under 256 MB while processing the 1 GB fixture, regardless of input size. Measured by: the platform’s process-accounting tool — /usr/bin/time -l on macOS, /usr/bin/time -v on GNU/Linux — with maximum resident set size recorded for the 1 GB fixture and for a 4 GB one. If the second number is materially larger than the first, you are not streaming.

NFR-SEC-01 — Secrets and the filesystem

Priority: Must Requirement: No credential or token appears in the repository at any commit in its history, and the tool writes to no path other than standard output, standard error, the path given by --out, and the rejects path derived from it. Measured by: a secret-scanning step in continuous integration over full history; plus a test that runs the tool inside a temporary directory tree and asserts that no file was created outside the declared output paths. Log files routinely contain addresses, tokens in query strings, and personal data — a tool that writes them somewhere unexpected is a tool that leaks.

NFR-ACC-01 — Readable without color

Priority: Must Requirement: Severity shall never be signaled by color alone — every record is prefixed with ERROR, WARN, or INFO; the tool shall honor the NO_COLOR convention and a --no-color flag; and default output shall fit within 100 columns. Measured by: a golden-file test asserting that output contains no ANSI escape sequences when NO_COLOR is set; a manual read of the report through a screen reader or piped to a file. A command-line tool has neighbors too.

NFR-REL-01 — Corrupt input never halts the run

Priority: Must Requirement: No malformed, truncated, or mis-encoded line shall terminate a run; the tool completes, reports, and exits with a documented code in every case. Measured by: a golden-file test over a deliberately corrupt fixture containing a truncated final line, an invalid UTF-8 byte sequence, and an embedded newline inside a quoted field; the run must exit with the documented code and produce a non-empty rejects file.

NFR-MAINT-01 — Clean clone to green tests

Priority: Should Requirement: A stranger with a clean machine shall get from git clone to a passing test run in under 10 minutes, using only script/setup and script/test. Measured by: the clean-machine test in Week 7, performed by a person who is not you, timed, with every place they got stuck written down.

The genuinely hard part

Defining “anomalous” so that you can defend it in Week 8. Three standard deviations above a trailing sixty-minute mean sounds rigorous, and it is arbitrary in at least four places, all of which a reviewer will find. Error counts per minute are not normally distributed, so a standard-deviation rule is a heuristic wearing statistical clothes. A large burst inside the baseline window inflates the mean and the standard deviation, which hides the next burst — the failure mode where an outage suppresses its own alarm. A quiet period drives the standard deviation toward zero, so two errors at 4 a.m. trip a threshold and Priya stops trusting you. And the first sixty minutes of any file have no baseline at all, which is why FR-DETECT-01 has a warm-up criterion. You do not have to solve all of that. You have to know all of that, choose deliberately, and write the choice down with its known weaknesses — which is a better Week-8 answer than a rule that happens to work on your one fixture.

The second eater is the seam between the parser and the detector, and it fails silently, which makes it the most dangerous defect in the project. If the parser represents an unparseable line as “nothing” and the detector assumes every minute it does not see had zero errors, then a file that is 90% corrupt produces a confident, clean, entirely wrong report saying nothing happened. Decide, in the technical specification, what a rejected line means to the detector, and write the test that proves it — feed the detector a fixture that is half rejects and assert the report says so.

A stack that fits

A language you already know, with as close to no dependencies as you can manage. Python’s standard library alone will do everything in the Must list — file iteration, a JSON parser, argparse, statistics — and if you are quicker in Go or Rust, both give you a single distributable binary, which makes Week 7’s packaging nearly free. What matters is not which one; it is that you can stream a file line by line without loading it, and that your dependency list is short enough for a stranger to install in ten minutes.

Resist two temptations specifically. Do not reach for a data-frame library because the word “analysis” appeared: loading a 1 GB log into a data frame is how you fail NFR-PERF-02 in a single import statement. And do not put a web interface on this. Packaging is deployment for a command-line project, and spending eight hours putting TraceLens behind a web server to look impressive is a real, gradeable mistake — Week 7 asks for a tagged release with an installable artifact and instructions somebody else verified, not a hosted service.

Novelty load: this brief is unusually forgiving because there is no framework in it. If you have written Python and used argparse, your novelty load is close to zero and every hour you budget is an honest hour. If you choose it as your excuse to learn Rust, your novelty load is one large thing and you should cut a Should before Week 5 rather than after. Either way, the choice goes in an architecture decision record in Week 3, with the alternatives you actually considered and why you rejected them. This is a recommendation, not a verdict.

What to cut first if Week 4 says you are behind

  1. Cut FR-PARSE-02, format auto-detection. Saves 6 hours. Require --format and say so in the README. The tool is slightly worse and completely honest.
  2. Cut FR-REPORT-02, the JSON output. Saves 5 hours including the schema and its golden tests. Human output was always the primary path.
  3. Cut FR-CLI-02, the time-range filter. Saves 4 hours of date parsing, timezone handling, and the boundary tests that come with it. Operators can pre-filter with tools they already have.
  4. Narrow NFR-PERF-01 from 1 GB to 200 MB, and say so in writing. Saves roughly 4 hours of fixture generation and profiling. Do not silently stop measuring — restate the threshold with its reason in the change log.

If you are ahead

Add a compressed-input path and a directory mode: accept gzip-compressed logs transparently and accept a directory of rotated files, processing them in timestamp order as one continuous stream so a baseline survives a file boundary. It sounds small and it is not — rotation boundaries are where real detectors break, and handling them is the most operationally credible thing in the whole project. Budget 10 hours, and only after the Must list is tested, documented, and packaged.


Brief 3 — SlotCheck

One line: A web application that lets a campus organization post shifts, lets members claim them, refuses to double-book anybody, and reminds them before they are due.

The person who has this problem: Nadia, 20, volunteer coordinator for a forty-member campus service club. She runs sign-ups in a shared spreadsheet that anyone can overwrite and a group chat that nobody reads past the second message. Her Sunday nights go to two chores: finding out which slots nobody claimed, and texting the three people who signed up for two things at the same time. Nadia is a sketch. Your version of her is a real coordinator — of a club, a lab, a tutoring center, an intramural league, a church nursery rota — whom you talk to in Week 1 and quote in your requirements document. This brief is worthless without one, because you will not know what “conflict” means at your institution until somebody tells you.

Why it fits 160 hours: A vertical slice costs seven to nine hours in a stack you already know; these four Musts price out bottom-up at 9, 6, 10, and 4 — the claim path above the range because of the concurrency work, the roster view below it — for about 29 hours of feature construction, which is the whole of what the 40-hour build-and-verify window buys after 11 hours of tests and defect work. Add about 8 hours of walking skeleton and continuous integration in Week 3 and 6 hours of deployment in Week 7. This is the largest of the first four briefs and it has the least slack, because sign-in is expensive and the conflict rule is harder than it looks. Take the two-thirds rule seriously here or it will take you seriously in Week 6.

The problem

The spreadsheet works until it does not. Forty people have the link, which means forty people can edit any cell, which means the version Nadia looks at on Sunday is not the version she published on Thursday. Somebody sorted it. Somebody deleted a row to “clean it up.” Two people typed their names into the same cell within a minute of each other and one of them silently won, and neither of them knows which. The failure is not that the spreadsheet is a bad tool; it is that it has no idea what a slot is, so it cannot enforce anything at all.

The chat is worse, and it is where the coordination cost actually lands. A shift opens up, Nadia posts it, four people respond over six hours, and now she is the merge conflict resolver: she has to read a conversation to determine who claimed what, and then remember it, and then transcribe it. When somebody drops out on Friday afternoon there is no mechanism at all — she scrolls, she guesses who is free, she sends direct messages. Every one of those minutes exists only because no system anywhere knows who is committed to what.

Then there is the conflict problem, which is the one that makes this a capstone instead of a form. People sign up for overlapping things constantly, not because they are careless but because the overlap is invisible at the moment they commit. They see one event’s page. They do not see their own week. And a coordinator only finds out on the day, which is exactly when it cannot be fixed. What Nadia actually needs is not a sign-up sheet — it is a system that knows enough about time to say no.

Minimum viable scope (the Must list)

  • Get people in without building an identity system. Passwordless sign-in by emailed link, plus membership in exactly one organization. (~9 h)
  • Post an event with slots. A coordinator creates an event and the slots under it — start, end, capacity — and publishes it to the organization. (~6 h)
  • Claim and release a slot, safely. A member claims a slot; the system enforces capacity and refuses a claim that overlaps in time with something that member already holds. Releasing puts the seat back. (~10 h)
  • See the roster and fix it. A coordinator sees every slot with who holds it and what is unfilled, and can remove or reassign anyone. (~4 h)

Reminders are in this brief and they are a Should, not a Must, for one reason: delivery is the part you do not control. Build the reminder as an outbox row first — a record that says who should be told what and when — and render it in the interface. Then, and only then, wire an actual transport to drain the outbox. Done in that order, the feature is demonstrable in Week 8 even if mail delivery fails on stage, and cuttable in zero hours if Week 4 goes badly.

Four. Recurring events, waitlists, shift swapping between members, membership in more than one organization, and calendar export or subscription are not in the release.

Starter functional requirements

Ten drafted requirements. They describe a club I invented; the moment you have talked to a real coordinator, at least three of them will be wrong for your organization. Finding out which three is what Week 2 is for.

FR-ACC-01 — Passwordless sign-in

Priority: Must Requirement: A person shall be able to sign in by requesting a single-use link sent to their email address, which establishes a session lasting thirty days. Rationale: This project needs identity — you cannot detect a conflict without knowing whose calendar it is, and you cannot send a reminder without an address. It does not need passwords, and every password you store is a liability you chose to accept. A single-use link is one flow instead of five. Acceptance criteria:

  • Given a valid email address, when a person requests a link, then a single-use token is created with an expiry no longer than fifteen minutes and the interface states that a link was sent without revealing whether that address is registered.
  • Given a link that has already been used or has expired, when a person opens it, then they are refused with a generic message and offered a new link.
  • Given a successful sign-in, when the person returns within thirty days, then no new link is required.

FR-ACC-02 — Organization membership and roles

Priority: Must Requirement: A signed-in person shall belong to exactly one organization, in one of two roles — coordinator or member — and shall see only that organization’s events. Rationale: Two roles is the smallest model that supports the actual asymmetry: somebody creates the shifts and everybody else claims them. Multiple organizations per person is a data-model change disguised as a small feature. Acceptance criteria:

  • Given a member of organization A, when they request any page belonging to organization B, then the server refuses the request regardless of what the interface offered them.
  • Given a member (not a coordinator), when they attempt to create or delete a slot, then the server refuses with a permission error and no slot is created.

FR-EVT-01 — Create an event with slots

Priority: Must Requirement: A coordinator shall be able to create an event with a name, a date, and one or more slots, each with a start time, an end time, and a capacity of at least one. Rationale: The slot is the unit the whole system reasons about. Getting its shape right — bounded in time, bounded in headcount — is what makes conflict detection and capacity enforcement possible at all. Acceptance criteria:

  • Given a coordinator submitting an event with three slots, when they publish it, then all three appear to every member of the organization with their times and remaining capacity.
  • Given a slot whose end time is not after its start time, when the coordinator submits, then the system rejects the slot and names the field.
  • Given an event with no slots, when the coordinator publishes, then the system rejects the submission and states that at least one slot is required.

FR-SGN-01 — Claim a slot

Priority: Must Requirement: A signed-in member shall be able to claim a published slot that has remaining capacity, which reduces that slot’s remaining capacity by one. Rationale: This is the transaction the entire project exists to make trustworthy, and it is the one place where being approximately right is the same as being wrong. Acceptance criteria:

  • Given a slot with capacity 2 and one existing claim, when a member claims it, then the claim succeeds and remaining capacity reads 0 for every viewer.
  • Given a slot with no remaining capacity, when a member claims it, then the claim is refused with a message stating the slot is full, and no claim row is created.
  • Given twenty members submitting a claim on a one-seat slot at the same moment, when the requests are processed, then exactly one succeeds and nineteen are refused with the full-slot message.

FR-SGN-02 — Refuse an overlapping claim

Priority: Must Requirement: The system shall refuse a claim whose slot overlaps in time with any slot the same member already holds, stating which existing commitment it conflicts with. Rationale: This is the requirement Nadia’s spreadsheet cannot express and the reason her Sunday nights are spent on text messages. It is also where the project’s difficulty lives, because “overlaps” is a decision, not a fact. Acceptance criteria:

  • Given a member holding a slot from 14:00 to 16:00, when they claim a slot from 15:00 to 17:00 on the same date, then the claim is refused and the message names the 14:00 slot.
  • Given a member holding a slot ending at 16:00, when they claim a slot starting at exactly 16:00, then the claim succeeds — touching boundaries do not overlap, and this rule is stated in the specification.
  • Given a member holding a slot on one date, when they claim a slot at the same clock time on a different date, then the claim succeeds.

FR-SGN-03 — Release a claim

Priority: Must Requirement: A member shall be able to release a slot they hold, which returns one unit of capacity to that slot and removes their name from the roster. Rationale: People drop out. A system that makes dropping out invisible pushes the whole problem back into the group chat, which is where it started. Acceptance criteria:

  • Given a member holding a slot on a full event, when they release it, then remaining capacity increases by one and the slot becomes claimable by others.
  • Given a slot whose start time has already passed, when a member attempts to release it, then the release is refused and the message states that past slots cannot be released.

FR-ROS-01 — Roster view

Priority: Must Requirement: A coordinator shall be able to view, for one event, every slot with the members holding it and the number of unfilled seats, ordered by start time. Rationale: This is the screen Nadia opens on Sunday night, and the number she is looking for is the unfilled count. If the roster does not surface that in one glance, it has not replaced the spreadsheet. Acceptance criteria:

  • Given an event with four slots of which one is unfilled, when a coordinator opens the roster, then the unfilled slot is visually distinct and the header states the total number of unfilled seats.
  • Given an event with no claims at all, when a coordinator opens the roster, then every slot is listed with an explicit empty state rather than an empty page.

FR-ROS-02 — Coordinator override

Priority: Should Requirement: A coordinator shall be able to remove a member from a slot, or assign a member to a slot with remaining capacity, and the affected member shall see the change on their next visit. Rationale: Real coordination includes fixing things by hand. Without an override the coordinator’s only tool is asking the member to do it, which is the group-chat workflow you were replacing. Acceptance criteria:

  • Given a member holding a slot, when a coordinator removes them, then capacity returns and an audit row records who made the change and when.
  • Given a coordinator assigning a member to a slot that would overlap a commitment that member already holds, when they confirm, then the system warns, names the conflict, and requires an explicit confirmation before proceeding.

FR-NOT-01 — Reminder before a shift

Priority: Should Requirement: The system shall record, for every claimed slot, one reminder to be delivered to the holder at a coordinator-configured interval before the slot’s start time, and shall deliver each reminder at most once. Rationale: “Nobody showed up” is usually “nobody remembered.” Building the reminder as a durable outbox row rather than a fire-and-forget send is what makes at-most-once possible and what makes the feature demonstrable when the transport is down. Acceptance criteria:

  • Given a claimed slot and a twenty-four-hour reminder interval, when the reminder job runs, then exactly one pending reminder exists for that member and slot.
  • Given the reminder job is run three times in the same hour, when it completes, then the total number of reminders marked delivered for that slot is still one.
  • Given a member releases a slot before the reminder is delivered, when the job runs, then no reminder is delivered for that slot.

FR-EVT-02 — Recurring events

Priority: Won’t (this release) Requirement: A coordinator shall be able to create an event that repeats weekly until a stated end date. Rationale: Recorded so the decision is visible and so you are not tempted by it in Week 6. Recurrence is not a loop that creates rows — it is exceptions, edits that apply to one occurrence or all of them, and daylight-saving transitions that move a 9 a.m. shift by an hour twice a year. Twenty hours minimum. Coordinators can create four events. Revisit after a v1.0 ships.

Starter non-functional requirements

NFR-REL-01 — Capacity is never exceeded under concurrency

Priority: Must Requirement: No slot shall ever hold more claims than its capacity, under any pattern of concurrent requests. Measured by: an automated test that fires 20 simultaneous claim requests at a one-seat slot and asserts exactly one success, nineteen refusals with the documented message, and a database row count of exactly one; run in continuous integration on every commit. This is the single most important line in the brief — if it is enforced only by an application-level check-then-insert, this test will eventually fail, and it should.

NFR-PERF-01 — Roster render time

Priority: Must Requirement: The roster view renders at a 95th percentile under 1.5 seconds with a seeded organization of 40 members and 60 slots, on a throttled Fast 3G profile with a cold cache. Measured by: twenty loads in browser developer tools with throttling applied; the p95 recorded in the measurements log in Week 6 and again in Week 8.

NFR-SEC-01 — Server-side authorization on every write

Priority: Must Requirement: Every write endpoint shall verify, on the server, that the caller belongs to the target organization and holds the role the operation requires. No credential or token shall appear in the repository at any commit in its history. Measured by: one negative test per write endpoint asserting a 403 for a caller in another organization and, for coordinator-only operations, for a plain member; plus a secret-scanning step in continuous integration over full history. Hiding a button in the interface is not authorization — it is decoration over an open door.

NFR-PRIV-01 — Member data comes back out

Priority: Must Requirement: The system shall hold, for each member, only an email address, a display name, and their claims; and a member shall be able to request deletion, after which no row anywhere associates their email address with any past claim. Measured by: a data inventory table in docs/requirements.md listing every element, where it lives, how long it is kept, and how a member gets rid of it; plus a delete-and-query test that asserts zero matching rows across every table, including the reminder outbox and the audit log.

NFR-ACC-01 — Keyboard, contrast, and errors in words

Priority: Must Requirement: Every interactive control shall be reachable and operable by keyboard alone with a visible focus indicator; body text shall meet a contrast ratio of at least 4.5:1; every form input shall have a programmatically associated label; and a refused claim shall be announced in text, never by color alone. Measured by: unplug the mouse and complete claim, release, and roster review; run a contrast checker over every text color pair; set the display to grayscale and confirm a refused claim is still comprehensible.

NFR-REL-02 — The deployed system stays up

Priority: Should Requirement: The deployed application answers a health check successfully on at least 13 of 14 consecutive daily checks across the final two weeks. Measured by: a scheduled check appending pass or fail with a timestamp to the measurements log. If your hosting idles the service after inactivity, say so in the specification and measure what you actually have rather than claiming an availability you cannot deliver.

The genuinely hard part

Two people claiming the last seat at the same moment, and the fact that the obvious implementation is wrong in a way that passes every test you would think to write. Read the remaining capacity, see that it is one, insert the claim: two requests can both read “one” before either inserts, and now a one-seat slot has two people in it. This is a classic race and it will not show up in manual testing, because you cannot click twice at the same instant. It shows up on the day of the event when two people arrive for one shift. The fix is not cleverness in application code — it is letting the database enforce the invariant, inside one transaction, with a constraint or a lock, and then writing NFR-REL-01’s concurrency test to prove it. Do that in Week 5, on the walking skeleton, not in Week 6 on a finished feature.

And “overlap” is a policy question that engineers keep trying to answer with arithmetic. Do back-to-back slots at the same location conflict? Almost certainly not. Across campus, with ten minutes between them? Nadia has an opinion and it is not in any specification you can derive. Does a member who is waitlisted for one thing conflict with claiming another? Does a coordinator’s override get to break the rule? What happens on the Sunday when clocks change and a 2:30 a.m. slot exists twice or not at all? You will not answer all of these, and you should not try. Pick the three that your real coordinator says happen, write them into FR-SGN-02’s acceptance criteria as concrete times, and put the rest in the out-of-scope list with a reason. A conflict rule you can state in one sentence and defend beats a clever one nobody can predict.

A stack that fits

Whatever web stack you have already shipped something in, over a relational database with real transactions. The technology choice barely matters here; the database choice does, and specifically it must be one where you can express “insert this claim only if capacity remains” as a single atomic operation — a unique constraint plus a serialized transaction, a conditional update returning a row count, or a row-level lock. PostgreSQL gives you all three and one managed instance for the deployed system. SQLite is entirely legitimate for forty users and makes the handoff trivially reproducible, provided you understand its locking model and write the test that proves your understanding. Either way the enforcement lives in the database, not in an if statement.

For email, choose any transport you can obtain and verify yourself, and design so that you are not married to it: the application writes an outbox row, and a small dispatcher drains it. Free tiers, sending limits, and terms for transactional mail vary by provider and change without notice — as of 2026 you should read the current terms on the provider’s own page, write down what it said and the date you read it, and put that note in your assumptions section. Do not take a limit from a blog post, a classmate, or an assistant. Verify the whole path once, for real, in Week 1: send yourself one message and keep the transcript.

Novelty load is the risk that sinks this brief. Sign-in, a transactional claim path, an outbox with a scheduled job, and a deployed database is already four moving parts. If two of them are new to you, you are at the ceiling; if three are, cut the reminder before Week 5 rather than discovering it in Week 7. This is a recommendation to adopt or deviate from, and either way it becomes an architecture decision record in Week 3 — with the alternatives you weighed, the criterion that decided it, and the cost you accepted in writing.

What to cut first if Week 4 says you are behind

  1. Cut FR-NOT-01, the reminder, including the outbox and the scheduled job. Saves 8 hours. Everything you learn from the feature you already got from designing it; the roster still solves Nadia’s larger problem.
  2. Replace FR-ACC-01 with an organization join code and a typed display name. Saves 7 hours of email plumbing, token expiry, and single-use handling. The cost is real — anyone with the code can act as anyone — and it goes in the specification as an accepted risk with the reason, exactly as PantryPilot’s join code does. Note that this cut and cut 1 are the same cut: without email addresses there are no reminders.
  3. Cut FR-ROS-02, the coordinator override. Saves 5 hours including the audit trail and the conflict-warning path. Coordinators ask members to release; write that in the runbook.
  4. Cut FR-SGN-03, releasing a claim, down to “contact your coordinator.” Saves 3 hours. Do this one last and reluctantly — it makes the system meaningfully worse and it is the cut most likely to draw a question in Week 8. Have the answer ready.

If you are ahead

Add a personal week view: one screen showing every commitment a member holds across every event in the coming fourteen days, with the overlap rule visualized so they can see the collision the system would refuse before they attempt it. It turns a system that says no into a system that helps people plan, and it is the most credible answer to “what would you build next?” that this project has. Budget 9 hours including the date math and the empty state, and start it only when your Musts are tested, deployed, and documented.


Brief 4 — TrendDeck

One line: A small dashboard over one public dataset that answers a specific recurring question, with the ingest and transform reproducible from a single command.

The person who has this problem: Marisol, 41, operations director at a small nonprofit, who is asked the same question by her board every quarter and answers it by downloading a government spreadsheet, rebuilding the same pivot table, and pasting a chart into a slide. She has done this eleven times. The chart is slightly different every time and she is not entirely sure which version was right. Your Marisol might be a beat reporter at a local paper, a city-council staffer, a coach with a season of statistics, or a graduate student who re-runs the same aggregation every month — but you must have one, and you must be able to write down the sentence they say. This is the brief with the weakest natural user, and a dashboard without a question is a screensaver with axes.

Why it fits 160 hours: A vertical slice costs seven to nine hours in a stack you already know; these four Musts price out bottom-up at 7, 9, 8, and 6 — about 30 hours of feature construction against the 40-hour build-and-verify window, with roughly 10 hours reserved for tests and defects. Add about 8 hours of skeleton and continuous integration in Week 3 and 6 hours of deployment in Week 7 — and be warned that this brief hides its cost in a place the hour table does not show: the data will be dirtier than you estimated, and the transform is where the schedule goes.

The problem

The answer to Marisol’s board question exists. It is sitting in a public dataset that somebody’s agency publishes every quarter, and it has been sitting there the whole time. What does not exist is the path from that file to the number — because the file is a hundred thousand rows wide of things nobody at the nonprofit needs, with column names that changed in 2021, three date formats, and a category that was split into two categories partway through the period Marisol cares about. Every quarter she rebuilds that path by hand, in a spreadsheet, from memory.

That is the actual problem, and notice that it is not “there is no chart.” It is that the analysis is not reproducible. Nobody, including Marisol, can say exactly which rows were included in the number she showed the board in March, because the steps lived in a sequence of clicks that left no record. When a board member asks “is that up from last year?” the honest answer is “probably,” and that is a bad position for someone whose funding depends on the number.

There is a third piece, and it is the one that makes the project engineering rather than data entry. The source changes underneath you. A column gets renamed. A code list gains a value. A quarter arrives with a two-week delay and a footnote. Any pipeline that assumes last quarter’s shape will either crash — which is fine, it is loud — or quietly produce a plausible wrong number, which is not fine at all, because a plausible wrong number is indistinguishable from a right one on a chart.

Minimum viable scope (the Must list)

  • Reproducible ingest. One command pulls a pinned local copy of the raw data into a local store, reports the row count, and can be run twice without duplicating anything. (~7 h)
  • A transform that validates and refuses. One command turns raw rows into a documented analytical table, running checks — row-count band, null rates, unknown category values, date parseability — and exiting non-zero without writing anything if any check fails. (~9 h)
  • Three charts that answer the question. Not “some visualizations”: three specific views that answer the sentence your real user says. (~8 h)
  • One dimension filter and one time-range filter that every chart respects, with the current selection in the URL so a state can be shared. (~6 h)

Before any of this, in Week 1, run the Get gate for real: download the actual file, open it, record its size, its row count, its license, its update cadence, and the date you checked. I am not promising you that any particular dataset exists, is free, is currently published, or permits redistribution — availability and terms change without notice and vary by jurisdiction. As of 2026 many governments, transit agencies, and scientific archives publish open data under permissive terms, and many do not; verify the current terms on the publisher’s own page, and write the URL and the date into your requirements document. A brief that depends on a file you have only read about is not a project. It is a wish with a Gantt chart.

Four. Scheduled refresh from the live source, viewer accounts, saved or named views, alerting when a measure crosses a threshold, and a second dataset are not in the release.

Starter functional requirements

Nine drafted requirements, written to be dataset-agnostic on purpose. Replace every bracketed generality with the actual column, the actual category, and the actual number from your file — a requirement you cannot make concrete is a requirement you have not yet understood.

FR-ING-01 — Pinned, reproducible ingest

Priority: Must Requirement: script/ingest shall load a locally pinned copy of the raw dataset into the local store, reporting the source file’s name, its checksum, and the number of rows loaded. Rationale: Pinning a local copy — rather than downloading at run time — is what makes the pipeline reproducible in Week 7 when a stranger runs it, and what stops the publisher’s next release from silently changing your demo the night before it. Acceptance criteria:

  • Given the pinned raw file, when script/ingest runs on a clean machine, then the raw table exists with the documented row count and the checksum printed matches the one recorded in the README.
  • Given the raw file is missing or its checksum does not match, when script/ingest runs, then it exits non-zero naming the expected and actual values, and writes nothing to the store.

FR-ING-02 — Ingest is idempotent

Priority: Must Requirement: Running script/ingest twice in succession shall leave the store in the same state as running it once. Rationale: The most common quiet defect in a student pipeline is a re-run that doubles every row, which then doubles every total on the dashboard while every chart still looks entirely reasonable. Acceptance criteria:

  • Given a store already populated, when script/ingest runs again, then the row count is unchanged and the command reports that it replaced or skipped rather than appended.
  • Given a source file that changed since the last run, when script/ingest runs, then the affected rows are updated in place, the row count reflects only genuinely new records, and the run reports how many were added, updated, and skipped.

FR-XFM-01 — The analytical table

Priority: Must Requirement: script/transform shall produce a documented analytical table with one row per [unit of analysis] per [time period], with every column’s meaning, type, and units recorded in the data dictionary in the repository. Rationale: The analytical table is the contract between the messy half of the project and the visible half. Once it exists and is documented, the dashboard is straightforward; before it exists, nothing downstream can be trusted. Acceptance criteria:

  • Given the raw table, when script/transform runs, then the analytical table contains exactly one row per unit per period with no duplicates, verified by a uniqueness assertion in the test suite.
  • Given the data dictionary, when a stranger reads it, then every column in the analytical table appears with its type, its units, and one sentence saying what it means.

FR-XFM-02 — Validation that fails loudly

Priority: Must Requirement: script/transform shall run a documented set of validation checks — total row count inside an expected band, null rate per column below a stated threshold, every category value present in the committed code list, every date parseable — and shall exit non-zero without writing the analytical table if any check fails. Rationale: This is the requirement that separates a pipeline from a script. A wrong chart is worse than a missing chart, because somebody will act on the wrong one. Acceptance criteria:

  • Given a fixture with a category value absent from the code list, when script/transform runs, then it exits non-zero, names the offending value and its row count, and the analytical table is not modified.
  • Given a fixture in which one column’s null rate exceeds its threshold, when script/transform runs, then it exits non-zero naming the column, the observed rate, and the threshold.
  • Given the clean fixture, when script/transform runs, then all checks pass, the table is written, and the check results are printed as a summary.

FR-XFM-03 — Agreement with a published figure

Priority: Must Requirement: For at least one [period], the total the system computes for [the headline measure] shall match the figure published by the source, within a tolerance stated in the specification, with the source URL and the date checked recorded in the requirements document. Rationale: You cannot tell a correct chart from an incorrect one by looking at it, and neither can your grader. An external oracle is the only cheap way to know your joins and filters are right. It is also the single most credible slide in your Week-8 presentation. Acceptance criteria:

  • Given the committed expected value for the chosen period, when the test suite runs, then the computed total is within the stated tolerance.
  • Given a computed total outside tolerance, when the test suite runs, then the test fails with both values printed and the deployment is blocked.

FR-VIZ-01 — The three views

Priority: Must Requirement: The dashboard shall present three views — a trend over time, a comparison across [the primary dimension], and a distribution or ranking — each with a title stating the question it answers, and each labeled with the source and the date of the underlying data. Rationale: Naming the question in the chart title is a discipline, not decoration: a chart whose question you cannot write in eight words is a chart that should not exist. Acceptance criteria:

  • Given the analytical table, when the dashboard loads, then all three views render with axis labels, units, and a source line naming the publisher and the data’s as-of date.
  • Given a filter selection that returns no rows, when the views render, then each shows an explicit empty state naming the filter that excluded everything, not an empty axis.

FR-VIZ-02 — Cross-filtering with shareable state

Priority: Must Requirement: A viewer shall be able to select one value of [the primary dimension] and a time range, and all three views shall re-render against the filtered data; the current selection shall be encoded in the URL. Rationale: The URL is what makes a dashboard usable by more than one person: Marisol sends her board a link, not a screenshot. It is also, in practice, the cheapest way to make your own manual testing repeatable. Acceptance criteria:

  • Given a dimension value and a time range are selected, when the viewer copies the URL and opens it in a new session, then the same three filtered views render.
  • Given a time range with no data, when it is selected, then all three views show the empty state and the filter controls remain usable.

FR-VIZ-03 — Download the filtered data

Priority: Should Requirement: A viewer shall be able to download the currently filtered analytical rows as a comma-separated file, with the same column names as the data dictionary. Rationale: The single most common thing a real user does with a dashboard is try to get the numbers out of it so they can put them in a slide. Refusing that pushes them back to the spreadsheet you were replacing. Acceptance criteria:

  • Given an active filter, when the viewer downloads, then the file contains exactly the rows behind the current views and a header row matching the data dictionary.
  • Given a filter combination that matches no rows, when the viewer downloads, then the file contains the header row alone and the interface states that the current filters select no records, rather than producing an empty file with no explanation.

FR-ING-03 — Scheduled refresh from the live source

Priority: Won’t (this release) Requirement: The system shall fetch new data from the publisher automatically on a schedule and refresh the dashboard without intervention. Rationale: Recorded so the decision is visible. Automated refresh means handling a publisher who is late, who changes a column, or who republishes a corrected file — which is schema-drift detection, alerting, and a rollback path, and that is twenty hours minimum. A documented manual command that a person runs each quarter is honest, cheap, and exactly what Marisol needs. Revisit after a v1.0 ships.

Starter non-functional requirements

NFR-PERF-01 — Filter responsiveness

Priority: Must Requirement: A change to either filter shall re-render all three views at a 95th percentile under 1.0 second, with the full analytical table loaded at its documented row count. Measured by: twenty filter changes timed in browser developer tools with the production build; the p95 recorded in the measurements log. State the row count next to the number, because the threshold is meaningless without it.

NFR-REL-01 — Bad data never reaches a chart

Priority: Must Requirement: No analytical table shall be published when any validation check in FR-XFM-02 fails, and no deployment shall proceed when the oracle test in FR-XFM-03 fails. Measured by: a continuous-integration job that injects a corrupted row into a fixture and asserts a non-zero exit with no write to the analytical table; plus the oracle test running as a required check before deploy.

NFR-SEC-01 — Secrets and credentials

Priority: Must Requirement: No credential, API key, or token shall appear in the repository at any commit in its history; if the source requires a key, it shall be read from the environment, with a placeholder example file committed and the real one ignored. Measured by: a secret-scanning step in continuous integration over full history, returning zero findings on the release commit.

NFR-LIC-01 — License and attribution obligation

Priority: Must Requirement: The dataset’s license, its attribution requirement, and any redistribution restriction shall be recorded in docs/requirements.md with the publisher’s URL and the date checked, and the required attribution shall appear on every page of the dashboard. Measured by: open the deployed page and read the footer; open the requirements document and confirm the URL and date are present. This is an obligation, not a quality attribute — you do not get to negotiate it, and “I redistributed a dataset I was not licensed to redistribute” is not a bug, it is a breach.

NFR-ACC-01 — Charts a person can actually read

Priority: Must Requirement: Every view shall have a keyboard-reachable text alternative containing the same numbers as a table; no series shall be distinguished by color alone; and all text shall meet a contrast ratio of at least 4.5:1. Measured by: set the display to grayscale and confirm every series is still identifiable by label, pattern, or direct annotation; unplug the mouse and reach every table; run a contrast checker over the chart palette as well as the body text, which is the pair students always forget.

NFR-MAINT-01 — Clean clone to a rendered dashboard

Priority: Should Requirement: A stranger with a clean machine shall get from git clone to a rendered dashboard in under 15 minutes, running only script/setup, script/ingest, script/transform, and the documented start command. Measured by: the clean-machine test in Week 7, performed by somebody who is not you, timed, with every stumble written down and turned into a README fix.

The genuinely hard part

The transform, and it is not close. Every hour you did not budget will be spent here, on things that are individually trivial and collectively enormous: a column renamed between releases, three date formats in one field, a text encoding that turns one region’s name into mojibake, numbers stored with thousands separators, a category that was split in 2021 so any trend crossing that boundary is wrong unless you map it, and a join that looks fine and quietly fans out because the key is not as unique as the documentation claims. None of this is intellectually hard. All of it is slow, and it arrives one surprise at a time, which is the worst possible shape for a schedule with eight weeks in it. Budget nine hours for the transform, expect fourteen, and buy the difference by cutting a chart rather than by cutting validation.

The second hard part is subtler and it is what will get asked about in Week 8: you have no way to know you are right. A chart cannot look wrong. If a filter silently drops the rows with a null region, the chart still renders, still trends, still looks professional, and is false. This is why FR-XFM-03 is a Must rather than a nicety — reproducing one number that somebody else already published is the only cheap oracle available to you, and without it you are shipping confident output with no verification behind it. Get that check passing in Week 5, on the walking skeleton, with one number. Everything after that is comparatively safe.

A stack that fits

Python for the pipeline, a file-backed analytical store for the data, and whatever renders charts fastest for you. The pipeline is a sequence of small, testable functions over rows, and Python’s data ecosystem is where you will find the fewest surprises. For the store, DuckDB or SQLite both give you a single file that lives beside the repository, restores in seconds, and makes the Week-7 handoff nearly trivial — no service to run, no credentials, no hosted database to pay for. Prefer the boring one you already know.

For the dashboard, you have two honest shapes. A static build — the transform emits a small JSON or CSV of pre-aggregated rows and a plain browser page charts it — deploys anywhere, costs nothing to host, and has no server to keep alive, which makes NFR-PERF-01 easy and NFR-REL-02-style availability a non-issue. A small server lets you filter over more data than a browser wants to hold, at the cost of a running process. Pick the static shape unless your analytical table genuinely will not fit in a browser, and if you pick it, say so in the architecture decision record with the row count that justified it.

Novelty load: a charting library you have never used counts as one new thing, and an analytical store you have never used counts as another. That is your whole budget, spent before you write a line of the transform. If you have never done either, drop to a plotting approach you already know — even server-rendered images with a table beneath them is a defensible choice for this brief, provided NFR-ACC-01 is honored. This is a recommendation to adopt or deviate from, and either way the decision, its alternatives, and its accepted cost go into an architecture decision record in Week 3.

What to cut first if Week 4 says you are behind

  1. Cut FR-VIZ-03, the filtered download. Saves 4 hours. Cheapest cut with the least user pain, and it is a one-line note in the runbook telling Marisol to ask you for a file.
  2. Cut one of the three views in FR-VIZ-01, keeping the trend and the comparison. Saves 5 hours including its empty state and its text alternative. Restate the requirement as two views with a dated change-log row; do not leave a requirement in the document promising three.
  3. Cut FR-VIZ-02’s dimension filter, keeping only the time range. Saves 5 hours of cross-filter state, URL encoding, and the interaction tests. The dashboard becomes a report, and a report that is right beats a dashboard that is late.
  4. Narrow the scope of the data to one region or one category, and say so in the title. Saves 4 to 8 unpredictable hours in the transform, because most category-mapping misery is proportional to how many categories you kept. Do this before cutting validation. Never cut validation.

If you are ahead

Add schema-drift detection on the next release of the data: fetch the publisher’s newest file into a staging table, compare its columns, types, and code lists against the pinned copy, and produce a report naming every difference — new columns, missing columns, changed types, unseen category values, row-count movement outside the expected band — without touching the published analytical table. It is the honest answer to the FR-ING-03 you marked Won’t, it demonstrates that you understand why automated refresh is dangerous rather than merely expensive, and it is the most professional thing in the brief. Budget 10 hours, and start it only when your Musts are validated, tested, deployed, and documented.


Brief 5 — Grounded (answers drawn only from a document collection you control)

One line: A question-answering service over a fixed set of documents you own, which answers only from those documents, always shows the passages it used, and visibly refuses when it cannot find support.

The person who has this problem: Dolores Amaya, the undergraduate advising coordinator for a department of about four hundred majors, who answers the same thirty policy questions every August out of a ninety-page handbook, a course catalog, and a folder of dated memos that occasionally contradict each other.

Why it fits 160 hours: The corpus is small, fixed, and already in your possession, so there is no data-acquisition risk — the thing that kills most applied-AI capstones in Week 6. Weeks 1–2 buy the charter and the requirements (40 h, no feature code). Week 3’s walking skeleton is one document in, one question out, one citation shown, through the one shell this project gets — a command-line entry point or a single web page — running in CI (about 10 h of that week’s 20). The 40 construction hours of Weeks 5–6 split across the four Musts: ingestion, chunking, and indexing 8; retrieval and the bounded answer 10; citations, abstention, fallback, and disclosure 6; the evaluation harness plus two scored runs 6 — 30 hours of Must construction against the thirty-hour budget, with the remaining 10 held back for tests, integration, and the defect log, the way Week 6 will actually demand. Weeks 7 and 8 are documentation, deployment, handoff, and delivery. Read that thirty as a ceiling and not a target with room above it: this Must set sits exactly on the line, so the first Should you adopt has to displace something.

The problem

Dolores is not short of information. The answer to “can I count the summer internship toward the practicum requirement” is in the handbook, on page 44, in a sentence with three conditions in it. The problem is that the handbook is a PDF, the catalog is a different PDF, and the memo that changed the rule in 2024 is a third file whose name is advising-updates-final-v2. Full-text search inside a PDF reader matches literal words, so a student searching “internship” gets eleven hits and reads the wrong one. In August, Dolores answers that question by email about thirty times.

What she wants is not a chatbot. She has watched a chatbot confidently tell a student something wrong about a graduation requirement, and a student acted on it, and the consequence was a semester. What she wants is a box she can type a question into that comes back with the paragraph — the actual text, from the actual document, with the page number — plus a short summary she can paste into an email. And when the documents do not answer the question, she wants to be told that, in those words, rather than handed a fluent paragraph invented out of nothing.

That reversal is the whole engineering interest of this project, and it is why the brief exists. The generation step is the easy part and it is not where your hours go. The hard, gradeable, defensible work is the machinery around it: provenance that survives chunking, a threshold below which the system refuses to speak, a defined behavior when the model is slow or down, and a frozen evaluation set that tells you honestly whether any of it works. That is Chapter 3’s five-part envelope, built for real.

Minimum viable scope (the Must list)

  • Ingest a controlled corpus — register documents you have the right to use, extract their text, chunk it, and index it with per-chunk provenance (document title, page or section). (~8 h)
  • Answer inside a stated envelope — a natural-language question in, a bounded answer out, drawn only from retrieved passages, within a latency bound. (~10 h)
  • Cite and abstain — every answer names the passages it used and can show their exact text; when nothing retrieves above threshold, the system refuses rather than guesses. (~6 h)
  • Score yourself against a frozen set — a committed evaluation set of thirty questions and a repeatable run that produces a pass/fail report. (~6 h)

Four. User accounts, conversational follow-up questions, documents supplied by readers rather than by the corpus operator, a feedback control that rates an answer, and any fine-tuning of the model are not in the release.

Before any of this, in Week 1, run the model gate for real. Two of those four Musts assume you can call something that generates text — and so do four of the drafted requirements below, all of them Musts: FR-ASK-01, FR-CIT-01, FR-ABS-01, and FR-FBK-01. This book does not promise you that you can. Hosted free allowances exist for some vendors, for some accounts, in some regions; they change without notice, and several vendors want a card on file before they will open the account at all. A small model running locally is a real zero-marginal-cost route, but it needs memory, disk, and patience that neither a Chromebook nor a two-core cloud development container has — and Path A, the browser-only path that Appendix A §A.1 makes the default, is exactly that machine. So make one real call in Week 1, from the environment you will actually build in, and write into your assumptions what it cost, what the limit was, and the date you checked. Note the second half of the gate too: a hosted continuous-integration runner cannot reach a model on your laptop (§A.9), so if local is your only route, everything CI exercises runs against the stub.

If the gate fails, take the retrieval-and-citation shape instead, and take it in Week 1 rather than Week 5. Ingestion, provenance, retrieval, citation, abstention, and the frozen evaluation set — the whole gradeable spine of this brief — need no generation at all. FR-DIS-01’s passages-only view stops being an override and becomes the primary output: a question in, the ranked passages out with document, page, and verbatim text, and an explicit “the corpus does not appear to cover this” when nothing clears the threshold. FR-ASK-01 and FR-FBK-01 are then withdrawn together with a dated change-log row, FR-CIT-01 and FR-ABS-01 are adapted to describe passages rather than an answer, NFR-PERF-01 is restated against retrieval alone (NFR-PERF-02’s 800 ms, not 10 seconds, because there is no model in the path), and NFR-QUAL-01 becomes “the expected source document is in the top three for at least 24 of the 30 questions.” That shape is strictly smaller than the one above — it removes the bounded answer and the whole fallback path — so it never costs you hours; spend what it returns on extraction quality and the evaluation set. It is a legitimate finished capstone, it is defensible in Week 8, and the version of it you decide on in Week 1 is worth three of the version you are forced into in Week 6.

Starter functional requirements

Ten drafted requirements. Dolores is a sketch, so every one of them needs a Source: line naming the real advising coordinator you find in Week 1 before it earns a place in docs/requirements.md.

FR-COR-01 — Register a document

Priority: Must Requirement: A corpus operator shall be able to register a document file in a supported format, after which its text is extracted, chunked, and indexed, with the document title and the page or section number recorded on every chunk. Rationale: Provenance has to be attached at ingestion. Bolted on afterward it is guesswork, and guessed citations are worse than none. Acceptance criteria:

  • Given a forty-page text-bearing PDF, when the operator registers it, then the document appears in the corpus inventory with a chunk count greater than zero, and a query matching a distinctive phrase on page 12 returns a chunk labelled page 12.
  • Given a file that yields no extractable text (an image-only scan), when the operator registers it, then the system rejects the file, states the file name and the reason, and leaves no chunks from that file in the index.

FR-COR-02 — Replace or remove a document

Priority: Should Requirement: A corpus operator shall be able to replace or remove a registered document, after which no chunk from the previous version is retrievable. Rationale: The 2024 memo supersedes the handbook page. A system that answers from a withdrawn document is worse than the PDF folder Dolores already has. Acceptance criteria:

  • Given a registered document whose chunks are retrievable, when the operator replaces it with a revised file, then queries return only chunks from the new version, and the inventory records the replacement date.
  • Given a replacement that fails extraction, when the operator submits it, then the previous version remains registered and retrievable, and the operator is told the replacement did not take effect.

FR-ASK-01 — Ask a question

Priority: Must if the Week-1 model gate passes; otherwise withdrawn together with FR-FBK-01, and the passages-only shape described above takes its place Requirement: A reader shall be able to submit a question of up to 300 characters and receive either one answer of at most 150 words composed only from retrieved passages, or an abstention (FR-ABS-01), within 10 seconds. Rationale: This is the envelope. Length, count, source restriction, and latency are the four things you can specify about a probabilistic output; the wording of the answer is not one of them. Acceptance criteria:

  • Given a corpus containing the practicum policy, when a reader asks whether a summer internship counts toward the practicum, then a response of at most 150 words is returned within 10 seconds and every factual clause in it is present in at least one cited passage.
  • Given a question over 300 characters, when the reader submits it, then the system rejects the submission, states the limit and the actual length, and preserves the text the reader typed.

FR-CIT-01 — Cite the passages used

Priority: Must Requirement: The system shall display with every generated answer between one and five citations, each naming the document title and the page or section, each expandable to the verbatim passage text that was supplied to the model. Rationale: Dolores’s real deliverable is the paragraph, not the summary. The citation is also your only mechanism for catching a wrong answer before a student acts on it. Acceptance criteria:

  • Given an answered question, when the reader expands a citation, then the verbatim passage text is shown with its document title and page, and that text is byte-identical to the chunk stored at ingestion.
  • Given a generation response that references no retrieved passage, when the system prepares to display it, then the answer is discarded and the abstention behavior of FR-ABS-01 is shown instead.

FR-ABS-01 — Abstain below the relevance threshold

Priority: Must Requirement: When no retrieved passage scores at or above the configured relevance threshold, the system shall not generate an answer, and shall instead state that the corpus does not appear to cover the question and list the three closest passages with their sources. Rationale: The refusal is the feature. A system that always answers is a system whose answers carry no information about whether the corpus contains the answer. Acceptance criteria:

  • Given a corpus with no content about parking permits, when a reader asks about parking permits, then no generated answer is shown, the abstention message appears, and three passages with sources are listed.
  • Given the retrieval component raises an error, when a reader asks anything, then the abstention message is shown with a distinct wording indicating a system fault, and the fault is recorded per FR-LOG-01.

FR-FBK-01 — Fall back when the model call fails

Priority: Must if FR-ASK-01 is adopted; otherwise Won’t (this release) — there is no generation call to fall back from Requirement: When the generation call exceeds 10 seconds, returns an empty body, or returns a response that fails schema validation, the system shall present the top five retrieved passages with a message stating that the answer service is unavailable, and shall record the failure. Rationale: Your model endpoint will be slow or unavailable at some point in eight weeks, and the highest-probability moment is the Week 8 demonstration. Undefined behavior under failure is a demo failure. Acceptance criteria:

  • Given the generation dependency is stubbed to sleep past the timeout, when a reader asks a question, then within 12 seconds the five passages and the unavailability message are shown and no partial answer text appears.
  • Given the generation dependency returns malformed output, when a reader asks a question, then the same fallback view is shown and the malformed response is recorded with its length and the first 200 characters.

FR-EVL-01 — Run the evaluation set

Priority: Must Requirement: An operator shall be able to run the committed evaluation set of 30 questions in a single command and receive a report giving, per question, whether the response satisfied the envelope, whether the expected source document was cited, and whether the system abstained; plus the aggregate counts. Rationale: This report is your acceptance criterion for the whole feature, your Week 4 evidence that you are on pace, and the only defensible answer to “how do you know it works?” in Week 8. Acceptance criteria:

  • Given the evaluation set at tests/eval/questions.csv, when the operator runs the evaluation command, then a per-question table and an aggregate summary are written to a timestamped report file within 10 minutes.
  • Given the generation dependency is unavailable for six of the thirty questions, when the run completes, then those six are reported as dependency failures, counted separately from content failures, and the run still exits with a report.

FR-EVL-02 — Freeze the evaluation set

Priority: Should Requirement: The evaluation set shall be committed to the repository before construction of the answer path begins, and any subsequent change to it shall be accompanied by a dated row in the document change log stating what changed and why. Rationale: Once you have seen the results, every failing case starts to look like a case that does not count. Locking the set in advance is the same discipline a researcher uses to lock an analysis plan, and for the same reason. Acceptance criteria:

  • Given the evaluation set was committed in Week 3, when a grader inspects the history, then the commit predates the first commit touching the answer path.
  • Given a question is edited in Week 6, when the change is committed, then the change log in docs/requirements.md contains a dated row naming the question and the reason.

FR-DIS-01 — Label the output and let the reader refuse it

Priority: Should Requirement: The system shall label every generated answer as machine-generated, and shall offer a control that replaces the answer with the retrieved passages alone. Rationale: Disclosure is the minimum honesty a probabilistic feature owes its user, and the override is what makes Dolores willing to use it in August, when she does not have time to be wrong. Acceptance criteria:

  • Given a displayed answer, when the reader activates the passages-only control, then the generated text is removed from view and the cited passages remain with their sources.
  • Given an abstention rather than an answer, when the view renders, then no machine-generated label is shown, because nothing was generated.

FR-LOG-01 — Record every query

Priority: Could Requirement: The system shall record for every query the question text, the identifiers of the retrieved chunks, the top relevance score, whether it answered or abstained, and the end-to-end latency. Rationale: This log is how you tune the threshold in FR-ABS-01 with evidence instead of vibes, and it is the raw material for the Week 8 measurements section. Acceptance criteria:

  • Given ten queries have been answered, when the operator inspects the log, then ten rows exist, each with all five fields populated.
  • Given a query that raised an unhandled error, when the operator inspects the log, then a row exists marking the error, with the latency recorded up to the point of failure.

Starter non-functional requirements

NFR-PERF-01 — End-to-end answer latency

Priority: Must Requirement: The p95 end-to-end latency from question submission to displayed answer shall be at most 10 seconds against a corpus of at least 500 chunks. Measured by: three consecutive runs of the 30-question evaluation set on the reference machine named in docs/architecture.md, p95 computed over all 90 responses and recorded in the measurements log each iteration.

NFR-PERF-02 — Retrieval latency

Priority: Should Requirement: Retrieval alone shall complete in at most 800 ms at p95 on the same corpus. Measured by: the same 90 responses, using the retrieval timing recorded by FR-LOG-01.

NFR-PRIV-01 — Corpus containment and secret hygiene

Priority: Must Requirement: Document text shall leave the machine only to the model endpoints named in docs/architecture.md; no corpus file shall be committed to the repository except the small sample under tests/fixtures/ whose license is recorded; and no credential shall appear at any commit in history. Measured by: a continuous-integration step that fails if any path outside tests/fixtures/ matching the corpus pattern is tracked by git, plus a secret-scanning step over full history returning zero findings, on every push.

NFR-ACC-01 — Keyboard and contrast

Priority: Must Requirement: The three core tasks — ask, read the answer, open a citation — shall be completable using the keyboard alone with a visible focus indicator at every step; body text shall meet a contrast ratio of at least 4.5:1; and the answer region shall be announced to a screen reader when it updates. Measured by: unplug the mouse and complete all three tasks, recording every place you got stuck; run a contrast checker over every text-and-background token pair; one screen-reader pass on the answer view. Results recorded in the test results document.

NFR-REL-01 — The failure paths are tested, not hoped for

Priority: Must Requirement: The abstention path and both fallback paths shall be covered by automated tests that stub the model dependency to time out, to return an empty body, and to return malformed output, and those tests shall pass in continuous integration on every push. Measured by: the three tests exist in tests/, are named for FR-ABS-01 and FR-FBK-01, and appear in the CI run log.

NFR-QUAL-01 — Grounding and the pass bar

Priority: Must Requirement: Across the 30-question evaluation set, at least 24 responses shall satisfy the full envelope, and zero responses shall cite a passage that does not contain the statement it is cited for. Measured by: you, scoring each response against the three-point written scale committed alongside the evaluation set, with the scored sheet committed after each run. The grounding count is a hard zero; the 24 is the negotiable number.

The genuinely hard part

Evaluation, and the fact that you are grading your own homework. There is no ground truth waiting for you — you have to write the thirty questions, decide what a correct answer looks like for each, commit that before you build, and then resist the overwhelming pull to revise the set once you see which ones fail. Everything about this project rewards moving that work to Week 6 and everything about the schedule punishes it. Build the evaluation harness in Week 3, before the answer path is good, when it will tell you embarrassing things. That is the point.

The second schedule-eater is calibration, and it has no clean answer. Your relevance threshold trades wrong answers against useless refusals, and every value you pick is defensible only against measurements you have actually taken. Set it too low and the system invents policy; set it too high and it abstains on questions the handbook plainly answers, and Dolores stops using it in four days. You will also discover that PDF text extraction is a genuine research problem in miniature — multi-column layouts, running headers that land in the middle of your chunks, tables that extract as word salad, and page numbers that do not match the printed ones. Budget real hours for extraction quality, or restrict your corpus to formats you have verified extract cleanly and say so in the assumptions section of docs/requirements.md.

A stack that fits

One language for the whole thing, and let it be the one you already know. Python is the conventional choice here because text extraction, chunking, and search have mature libraries and because the evaluation harness is a script, not an application. Store chunks in SQLite and lean on its full-text search for lexical retrieval — it ships in the standard library of several languages, needs no server, and makes your clean-machine test trivial. Put the model behind one interface with two implementations: the real one and a stub that returns canned responses, which is what makes NFR-REL-01 cheap instead of painful. Put the shell — command line or a single web page — last, and keep it thin.

On the model itself: you need something you can call, either a hosted endpoint or a small model running locally on hardware you have. Nothing in this book guarantees you a zero-cost route to either one — that is what the Week-1 model gate above is for, and it is why the fallback shape is written down before you start rather than improvised in Week 6. Do not write a price, a rate limit, or a context length into your specification from memory or from a blog post. Look it up on the vendor’s own page, write the number with the date you checked it into your assumptions, and re-check it before Week 7. The same goes for any embedding model. And notice the novelty load: if a vector store, an embedding pipeline, a web framework, and a model API are all new to you at once, that is four unknowns in a 160-hour budget, and four is too many. Start lexical. Add embeddings only if the evaluation set says lexical is not enough — which is a sentence you can only say because you built the evaluation set first.

This is a recommendation, not a verdict. Adopt it or deviate from it, and either way record the decision as docs/adr/0001-choose-the-stack.md in Week 3 with the alternatives you rejected and why.

What to cut first if Week 4 says you are behind

  1. Drop embeddings and hybrid retrieval; ship lexical search only. Saves about 10 hours and removes an entire dependency. Report the honest quality cost from your evaluation runs.
  2. Drop the web interface; ship a command-line tool. Saves about 8 hours. Dolores would rather have a working command she can run than a half-built page. This also retires the accessibility work on the answer view — say so, and re-scope NFR-ACC-01 to the terminal output rather than deleting it.
  3. Restrict ingestion to plain text and Markdown; convert the PDFs by hand once. Saves about 8 hours and most of the extraction misery. Record it as an assumption with a consequence.
  4. Shrink the evaluation set from 30 questions to 20, keeping the pass bar proportional. Saves about 4 hours of scoring. This is the last cut, not the first — a project with no evaluation set has no acceptance criteria at all.

Read what those four cuts do not do: none of them removes the model. They drop embeddings, a web page, PDF ingestion, and ten questions, and every one of them leaves you still needing something to call. The cut that removes the model is the passages-only shape in the scope section, and the week to take it is Week 1, off the model gate — not Week 5, off a billing page.

If you are ahead

Add a second retrieval strategy and run a genuine comparison: hybrid lexical-plus-embedding retrieval with a re-ranking step, scored against the same frozen evaluation set, with the per-question deltas reported and a written account of where it helped and where it made things worse. About 15 hours, and it converts a build into a small piece of evidence — which is a far better Week 8 talk than one more feature.


Brief 6 — Repo Health (a repository-readiness reporter)

One line: A command-line tool that reads a repository checkout and reports, with evidence for every claim, how ready that repository is for somebody else to own.

The person who has this problem: Dele Okonjo, the newest engineer on a four-person platform team, handed forty-one internal repositories in his second week and asked which ones the team can safely be on call for.

Why it fits 160 hours: There is no network, no account, no user interface framework, and no data to acquire — the entire input is files on disk and git history, which means your test fixtures are repositories you construct in a temporary directory from a script. That is an unusually low-risk profile for a capstone. Weeks 1–2 are the charter and requirements (40 h). Week 3 stands up the walking skeleton: one check, one repository, one JSON report, in CI, at about 8 h of that week’s 20. The 40 construction hours of Weeks 5–6 split roughly: the check engine and result model 9; the documented check set 10; the multi-repository run and ranking 5; the two renderers 6; hardening against adversarial repositories 4; and 6 for integration and the defects that always appear.

The problem

Dele’s first Monday looked like this. Repository one: a README that says make dev, and a Makefile with no dev target. Repository two: a continuous-integration workflow that has been failing since March and nobody noticed because nobody is on the notification list. Repository three: a tests/ directory with forty test files and no command anywhere that runs them. Repository four: ninety-odd percent of the lines last touched by an engineer who left in November. Repository five is fine. Dele cannot tell which is which without opening each one, and forty-one repositories at twenty minutes each is more than a week of his life.

His manager does not want a report. She wants a ranked list and a reason — “these three are the ones to fix first, and here is what is wrong with each” — and Dele has to be right, because he is going to spend the team’s sprint on whatever he names. That means every claim the tool makes has to be backed by something Dele can put on a screen: not “documentation: poor” but “README.md contains no section matching a run-instruction heading, and the only fenced command in it, make dev, does not correspond to a target in Makefile.”

The interesting part is that repository health is not a fact. It is a judgment, and the moment you attach a number to it you have made a set of value claims — that a missing license matters more than a stale dependency, that single-author concentration is a risk rather than a fact about a small team. That is not a reason to avoid the number. It is a reason to make the number explainable, to publish the weights, and to check them against human judgment. Chapter 4 calls this giving a quality word a metric. This project is that exercise, taken all the way to a shipped tool.

Minimum viable scope (the Must list)

  • Analyze one repository checkout and produce one result object in which every check carries a verdict, the evidence behind it, and the remediation for it. (~9 h)
  • A documented rubric of at least ten checks, each with a written definition, a stated limitation, and an honest unknown verdict for repositories where it does not apply. (~10 h)
  • Analyze a set of repositories and produce a ranked comparison Dele can hand to his manager. (~5 h)
  • Two renderings of one result: machine-readable output for automation, and a human report for a person deciding what to do on Monday. (~6 h)

Four. A hosted dashboard, a continuous-integration annotation mode, automatic remediation of anything the tool finds, historical trend analysis, and per-ecosystem deep inspection are not in the release.

The same architecture carries the other developer-tool shapes in this catalog’s family — a linter for a specific artifact (a requirements document, a container definition, an API description), or a continuous-integration helper that annotates changes. Whichever you pick, keep the machine-readable result as the primary artifact and every human view as a rendering of it. That one decision is what makes the second and third output modes nearly free.

Starter functional requirements

Ten drafted requirements. Dele is a sketch; every one of them needs a Source: line naming the real engineer, team, or repository set you are writing for before it earns a place in docs/requirements.md.

FR-RUN-01 — Analyze one repository

Priority: Must Requirement: A user shall be able to run the tool against a path containing a git repository checkout and receive a result containing every enabled check with a verdict of pass, fail, or unknown. Rationale: Everything else in the tool is a view over this result. Getting the result shape right in Week 3 is what makes Weeks 5 and 6 additive instead of structural. Acceptance criteria:

  • Given a fixture repository with a README containing run instructions and no LICENSE, when the tool is run against it, then the documentation check reports pass with the matched heading as evidence and the license check reports fail with the searched file names as evidence.
  • Given a path that is not a git repository, when the tool is run against it, then it exits non-zero with a message naming the path and the reason, and produces no partial result file.

FR-RUN-02 — Analyze a set of repositories

Priority: Must Requirement: A user shall be able to run the tool against a list of repository paths and receive a single result containing every repository, ranked by score, with per-repository detail retained. Rationale: Dele’s problem is not one repository, it is forty-one. A tool that only does one is a checklist with extra steps. Acceptance criteria:

  • Given a list of ten fixture repositories, when the tool is run, then the result contains ten entries ordered by score with ties broken by a documented, deterministic rule.
  • Given one of the ten paths is unreadable, when the run completes, then the other nine are reported normally, the tenth appears with an error entry naming the cause, and the process exits non-zero.

FR-CHK-01 — The check contract

Priority: Must Requirement: Every check shall emit an identifier, a verdict, the evidence it observed, a one-sentence remediation, and a stated limitation describing what the check cannot see. Rationale: The limitation field is what keeps the tool honest. A check that reports pass for “has tests” when it only looked for a directory named tests must say so, or Dele will read the pass as a guarantee and be wrong in front of his manager. Acceptance criteria:

  • Given any check runs, when its result is inspected, then all five fields are present and non-empty.
  • Given a check that cannot determine an answer — no recognized dependency manifest, for instance — when it runs, then it emits unknown with evidence naming what it looked for, not fail.

FR-CHK-02 — Handoff-readiness checks

Priority: Must Requirement: The tool shall check for the presence and minimum substance of a README containing run instructions, a license file, a change log, and a declared entry point, reporting each as a separate check. Rationale: These are the four artifacts that decide whether a stranger can get the repository running, which is the definition of handoff readiness this course uses and the thing Dele actually needs to know. Acceptance criteria:

  • Given a repository with a README whose body contains a fenced command block under a heading matching the configured run-instruction patterns, when the documentation check runs, then it reports pass with the matched heading and command as evidence.
  • Given a README that exists but contains fewer than the configured minimum number of words, when the check runs, then it reports fail with the word count as evidence and does not crash on an empty file.

FR-RPT-01 — Machine-readable report

Priority: Must Requirement: The tool shall write its complete result as a machine-readable document conforming to a schema committed in the repository, and shall exit with a documented status code indicating success, findings above the configured threshold, or tool error. Rationale: The schema is what lets the human report, any future continuous-integration mode, and your own tests all read the same object. The exit codes are what let it be used in a pipeline without parsing text. Acceptance criteria:

  • Given a completed run, when the output is validated against the committed schema, then validation passes and the document contains the tool version and the analysis timestamp.
  • Given the output path is not writable, when the run completes, then the tool exits with the documented tool-error code and states the path and the reason on standard error.

FR-RPT-02 — Human report

Priority: Must Requirement: The tool shall render a human report from the machine-readable result showing, per repository, the score, the failed checks ordered by weight, and the remediation for each. Rationale: This is the artifact Dele shows his manager on Monday. If reading it requires the tool’s author to be in the room, the tool has failed. Acceptance criteria:

  • Given a result covering ten repositories, when the human report is rendered, then the three lowest-scoring repositories appear first with their failed checks and remediation lines visible without scrolling past unrelated detail.
  • Given a result in which every check for a repository is unknown, when the report renders, then that repository is shown as unassessable with the reason, and is excluded from the ranking rather than scored as zero.

FR-CHK-03 — Verification checks

Priority: Should Requirement: The tool shall report whether a test directory or test-named files exist, whether a test command is declared in a recognized manifest or script location, and whether a continuous-integration workflow definition is present. Rationale: “Has tests” and “runs tests” are different facts and students conflate them constantly. So do repositories. Acceptance criteria:

  • Given a repository with test files, a declared test command, and a workflow file, when the checks run, then all three report pass with the specific paths as evidence.
  • Given a repository with test files but no declared command anywhere, when the checks run, then the presence check passes, the command check fails, and the remediation names the file the command should be declared in.

FR-CHK-04 — History checks

Priority: Should Requirement: The tool shall report, from git history, the proportion of tracked lines last modified by the single most frequent author, and the number of days since the most recent commit touching source files. Rationale: This is the closest an automated tool gets to Dele’s real question — if one person disappears, is this repository owned by anyone? It is a proxy, and the limitation field must say so. Acceptance criteria:

  • Given a fixture repository with a known authorship distribution, when the check runs, then the reported proportion matches the constructed value within one percentage point.
  • Given a repository with a single commit, or with no commits at all, when the check runs, then it reports unknown with the commit count as evidence and does not divide by zero.

FR-CFG-01 — Configuration

Priority: Should Requirement: A user shall be able to supply a configuration file that enables or disables individual checks and sets each check’s weight and thresholds, with documented defaults used when no file is supplied. Rationale: The weights are value judgments. Making them editable is how you avoid arguing about them and how you let a team encode its own standard. Acceptance criteria:

  • Given a configuration that disables the license check and doubles the documentation weight, when the tool runs, then the license check is absent from the result and the scores reflect the new weight.
  • Given a configuration naming an unknown check identifier, when the tool runs, then it exits with the tool-error code and names the unknown identifier and the valid identifiers.

FR-WVR-01 — Waivers with an expiry

Priority: Could Requirement: A repository shall be able to carry a committed waiver file recording an accepted exception per check, each with a reason, an owner, and an expiry date; waived checks are reported as waived rather than failed until the expiry date passes, after which they are reported as failed with the expired waiver as evidence. Rationale: A suppression with no expiry is a permanent lie that accumulates. The expiry is what turns “we know, it is fine” into a decision somebody has to renew. Acceptance criteria:

  • Given a waiver for the license check with a future expiry, when the tool runs, then the check reports waived, cites the reason and owner, and the score is computed with the check excluded.
  • Given a waiver whose expiry has passed, when the tool runs, then the check reports fail, and the evidence names the expiry date and the owner.

Starter non-functional requirements

NFR-PERF-01 — Analysis throughput

Priority: Must Requirement: The tool shall analyze a repository containing at least 20,000 commits and 5,000 tracked files in under 30 seconds, and a set of forty such repositories in under 8 minutes, on the reference machine named in docs/architecture.md. Measured by: timed runs over the fixture corpus, five runs each, median and worst case recorded in the measurements log each iteration.

NFR-SEC-01 — Read-only, and never execute the subject

Priority: Must Requirement: The tool shall not write to, modify, or execute any file inside a repository under analysis, and shall not execute any command declared by that repository. Measured by: an automated test that runs the tool against a fixture repository containing a hook, a build file, and a script that each write a sentinel file, then asserts no sentinel exists and the repository’s working tree is unchanged by checksum.

NFR-REL-01 — Determinism and hostile input

Priority: Must Requirement: Two runs over the same checkout shall produce identical machine-readable output apart from the timestamp field, and the tool shall complete without an unhandled exception against every repository in the adversarial fixture set. Measured by: a CI step that runs the tool twice and diffs the output with the timestamp removed, plus a suite over at least eight adversarial fixtures — empty repository, no commits, binary-only, symlink loop, submodule, file larger than 100 MB, filenames with newlines, and a detached HEAD.

NFR-ACC-01 — Output that does not depend on color

Priority: Must Requirement: Every verdict in terminal output shall be identifiable from text alone, color shall be suppressed when the NO_COLOR convention is in effect or when output is not a terminal, and any rendered HTML report shall meet a contrast ratio of at least 4.5:1 for body text. Measured by: capture piped output and confirm every verdict is textually distinguishable with no escape sequences present; run a contrast checker over every token pair used in the HTML renderer.

NFR-QUAL-01 — Agreement with human judgment

Priority: Should Requirement: The tool’s ranking of ten real repositories shall agree with the median ranking produced independently by three engineers on at least seven of the ten adjacent pairwise comparisons. Measured by: three raters rank the same ten repositories from the checkouts alone with no access to the tool; you record every disagreement and write a paragraph on each in the test results document. Disagreements are findings, not failures.

The genuinely hard part

The score is a value judgment wearing a number, and defending it is the project. Any weighting you choose is arguable, every check has a false-positive story, and the first time the tool tells a senior engineer that her repository scores 62 you will be asked to justify the 62. The engineering answer is to make the number fully decomposable — every point traceable to a check, every check traceable to evidence, every weight visible in a configuration file with a written rationale — and then to validate the ranking against people, which is what NFR-QUAL-01 exists for. That validation is the most interesting eight hours in this project and the first thing students cut. Do not cut it; cut a renderer instead.

The second thing that will eat the schedule is heterogeneity. Forty-one repositories in six languages means a check written against one ecosystem’s manifest reports unknown for most of the corpus, and a report that is eighty percent unknown tells Dele nothing. Deciding which checks are genuinely language-agnostic — documentation, license, change log, history, workflow presence — and which are ecosystem-specific, then designing unknown as a first-class verdict rather than an embarrassment, is architectural work that has to happen in Week 3. Get it wrong and you will be rewriting the result model in Week 6.

A stack that fits

Whatever language you can write a solid command-line tool in without looking things up. Python, Go, and Node are all fine here and the choice barely matters; what matters is that you shell out to git rather than binding a git library. Shelling out costs you a subprocess and buys you the entire feature set of a tool that is already installed everywhere you will run this, and it drops your novelty load by one. Design the machine-readable result first, commit its schema, and treat the terminal renderer and any HTML renderer as pure functions of it. Build fixtures with a script under script/ that constructs repositories in a temporary directory at test time — committing nested repositories inside your repository is a trap that will confuse both git and your grader.

Novelty-load warning: the temptation here is to pick a new language “because it is a good fit for CLI tools.” A tool you can write is worth more than a tool you can almost write in a language you are learning. If you do take the new language on purpose, say so in the charter as a stated learning objective with hours attached, and cut a check to pay for it.

Recommendation, not verdict. Adopt or deviate in Week 3 with docs/adr/0001-choose-the-stack.md, and record the git-subprocess decision as its own ADR if you go the other way.

What to cut first if Week 4 says you are behind

  1. Drop the HTML renderer; ship terminal output and the machine-readable document. Saves about 8 hours and loses very little, because Dele lives in a terminal.
  2. Drop multi-repository mode; ship single-repository plus a documented shell loop. Saves about 8 hours. Say plainly in the README that ranking across repositories is out of scope for this release, and put it on the Won’t list with a reason.
  3. Cut the check set from twelve to six, keeping the four handoff-readiness checks and two verification checks. Saves about 8 hours and makes the rubric easier to defend, not harder.
  4. Replace the human-validation study with a written weight rationale and a sensitivity table showing how the ranking moves when each weight is halved and doubled. Saves about 8 hours. This is genuinely the last cut — it costs you the most interesting claim you could make in Week 8.

If you are ahead

Add trend mode: re-run the analysis at a series of historical commits and report each repository’s score over time, so Dele can distinguish a repository that is bad and improving from one that is bad and rotting. About 15 hours, most of it in making the analysis cheap enough to run fifty times per repository and in handling checkouts that fail to build a working tree at old commits. It also gives you something genuinely new to say in the presentation.


Brief 7 — Own Audit (a secrets scanner for systems you administer)

One line: A scanner and triage workflow you point only at repositories and hosts you personally administer, which finds committed secrets, redacts everything it reports, and tracks every finding to a recorded decision.

The person who has this problem: Amara Okafor, the sole maintainer of a student organization’s code hosting account — fourteen repositories, three of them public — and the named administrator of the single virtual machine those repositories deploy to. She is the only person with the authority to audit any of it, and that authority is what makes this project legal.

Why it fits 160 hours: The scope is bounded by an artifact you write yourself: a scope file that lists exactly what may be scanned. There is no external service to integrate, no data to acquire, and no user interface. Weeks 1–2 are charter and requirements (40 h). Week 3’s skeleton scans one repository’s history for one rule and writes one redacted finding, in CI, at about 8 h of that week’s 20. The 40 construction hours of Weeks 5–6 split roughly: the scope file, inventory, and runner 7; history scanning and the rule set 11; redaction and the report 6; the triage and suppression store 6; the fixture corpus and the precision measurement 6; and 4 for integration and the defects that always appear.

The problem

Amara inherited the organization’s account from someone who graduated. She knows there is at least one API key in a commit somewhere, because she found one by accident in March while reading an old branch — a key that had been live for two years in a public repository. She rotated it. She has no idea whether there are others, and she has no way to find out except by reading fourteen repositories’ worth of history by hand.

Finding the key is not actually the hard problem, and this is the part that surprises students. A regular expression finds keys. What Amara does not have is a process: a record of what she looked at, what she found, what she decided about each one, and what she did about it. Without that, every scan starts from zero, the same forty false positives get re-examined every time, and after two runs she stops running it. A security tool that produces an unmanaged list of findings does not make a system safer; it makes a person tired.

Now the constraint that defines this brief, and it is not negotiable. This project is strictly defensive and strictly limited to systems you own or administer, or systems for which you hold written, dated authorization from the person who does. That authorization, if any, lives in your repository. You do not scan third-party repositories, you do not scan your employer’s systems without written permission from someone empowered to give it, and you do not publish findings about anybody else’s systems. You do not exploit anything. And you never send a discovered credential to the service it belongs to in order to see whether it still works — that is unauthorized use of someone else’s service, and the fact that you found the key does not make it yours. If you find a credential that belongs to someone else, stop, do not test it, do not commit it anywhere, and tell the owner. Your instructor is the person to ask before you do anything you are unsure about; the answer arrives faster than the consequences do.

Minimum viable scope (the Must list)

  • A scope file with an authority line per entry, and a tool that refuses to scan anything not listed in it. The refusal is a feature, and it is the one you demonstrate first. (~7 h)
  • Full-history secret detection across in-scope repositories, producing findings with file, commit, line, rule, and severity. (~11 h)
  • Redaction everywhere, so that no detected secret value appears in cleartext in any report, log, or pipeline output. (~6 h)
  • Triage that persists: every finding is confirmed, a false positive, or an accepted risk, with a reason, an owner, and a date, and that decision survives the next run. (~6 h)

Four. Vulnerability-advisory matching, automatic credential rotation, any form of exploitation or live credential testing, scanning of hosts or networks beyond the scope file, and a hosted service are not in the release.

The same skeleton carries the other two defensive shapes in this family. An authentication hardening lab replaces the scanner with a set of negative tests against your own deployed application — session fixation, missing server-side authorization on every write path, password reset token reuse — each one a requirement with a Given/When/Then. A log-based reporter replaces it with anomaly reporting over logs from a host you administer, which is the TraceLens shape from the main text. All three keep the scope file, the triage store, and the redaction rule unchanged.

Starter functional requirements

Ten drafted requirements. Amara is a sketch; every one of them needs a Source: line naming you, the systems you actually administer, and the date your authority over each was established, before it earns a place in docs/requirements.md.

FR-SCP-01 — Nothing is scanned that is not in scope

Priority: Must Requirement: The tool shall read a committed scope file listing every repository path and host in scope and shall refuse, with a non-zero exit, to scan any target not listed in it. Rationale: This is the control that makes the project defensible, and it is enforced by code rather than by intention. It is also the first thing to demonstrate in Week 8. Acceptance criteria:

  • Given a scope file listing three repositories, when the tool is run against one of them, then the scan proceeds and the report names the scope file and its checksum.
  • Given a target not listed, when the tool is run against it, then no file is read from that target, the tool exits non-zero, and the message states that the target is not in scope and names the scope file.

FR-SCP-02 — Every entry records its authority

Priority: Must Requirement: Each scope entry shall record who owns the target, the basis of your authority to scan it, and the date that authority was established; the tool shall refuse to scan any entry missing one of these fields. Rationale: “I own it” is a claim, and a claim you wrote down with a date is the difference between an audit and an incident. This field is also what makes the repository safe to show an employer. Acceptance criteria:

  • Given an entry with owner, basis, and date populated, when the tool validates the scope file, then validation passes and the three fields appear in the report header.
  • Given an entry missing the basis field, when the tool validates the scope file, then it exits non-zero naming the entry and the missing field, and scans nothing.

FR-SCR-01 — Scan full history

Priority: Must Requirement: The tool shall scan every blob reachable from every branch and tag of an in-scope repository, not only the current working tree, and shall report each finding with repository, commit, file path, line number, and rule identifier. Rationale: The dangerous key is almost never in the current tree. It was removed in a later commit by someone who believed that deleted it. Acceptance criteria:

  • Given a fixture repository where a key was added in commit A and removed in commit B, when the scan runs against the current tip, then the finding is reported with commit A and the file path as it existed then.
  • Given a repository containing a blob larger than the configured size limit, when the scan runs, then that blob is skipped, a notice naming the path and its size appears in the report, and the scan continues to completion.

FR-SCR-02 — Redact before reporting

Priority: Must Requirement: The tool shall never write a detected secret value in cleartext to any report, log, terminal output, or pipeline output; it shall emit instead a fingerprint consisting of the first four characters, the length, and a one-way hash of the value. Rationale: A findings report full of live credentials is a new vulnerability that you created, and it will end up in a pipeline log that is more widely readable than the repository was. Acceptance criteria:

  • Given a finding for a known planted value, when any output artifact is produced, then the fingerprint appears and the value does not, and two findings with the same value share a fingerprint.
  • Given a rule that captures a multi-line block, when the finding is emitted, then the entire captured block is fingerprinted rather than partially printed, and no line of it appears verbatim.

FR-TRI-01 — Triage a finding, once

Priority: Must Requirement: An auditor shall be able to record for any finding a disposition of confirmed, false positive, or accepted risk, together with a reason, an owner, and the date, stored in a committed file keyed by a stable finding identifier, and subsequent runs shall carry that disposition forward. Rationale: This is the feature that separates a tool Amara runs twice from a tool she runs monthly. It is also the hardest thing in the project to get right, because the identifier has to be stable across history rewrites. Acceptance criteria:

  • Given a finding triaged as a false positive with a reason, when the scan is run again, then the finding appears in the report as an already-triaged false positive with the original reason and date, and is excluded from the new-findings count.
  • Given a triaged finding whose underlying commit no longer exists after a history rewrite, when the scan runs, then the triage record is reported as orphaned with its original reason rather than silently dropped.

FR-RPT-01 — A report and an exit code

Priority: Must Requirement: The tool shall produce a report listing new findings, already-triaged findings, and orphaned triage records separately, and shall exit with a documented status code distinguishing “no new findings”, “new findings above the configured severity”, and “tool error”. Rationale: The three-way split is the whole point: Amara only ever reads the first section, and the exit codes are what let this run unattended later. Acceptance criteria:

  • Given two new findings and five triaged ones, when the run completes, then the report shows the two under new findings with full detail, summarizes the five, and exits with the new-findings code.
  • Given the triage file is malformed, when the run starts, then the tool exits with the tool-error code, names the line that failed to parse, and does not overwrite the triage file.

FR-SCR-03 — A rule set with stated severities and limitations

Priority: Should Requirement: The tool shall carry a documented rule set in which each rule states what it detects, its severity, and its known limitations, and a user shall be able to enable or disable individual rules by identifier. Rationale: Severity that is not written down gets invented at report time. Limitations that are not written down get read as guarantees. Acceptance criteria:

  • Given a rule disabled by identifier, when the scan runs, then no finding from that rule appears and the report lists the rule as disabled.
  • Given an unknown rule identifier in the configuration, when the tool starts, then it exits with the tool-error code and lists the valid identifiers.

FR-DEP-01 — Dependency inventory

Priority: Should Requirement: The tool shall produce, per in-scope repository, an inventory of declared dependencies with their declared version constraint and their locked version, and shall flag any dependency declared without a corresponding lock entry. Rationale: Drift between declaration and lock is where a reproducible build quietly stops being reproducible, and it is checkable with no external service. Acceptance criteria:

  • Given a repository with a manifest and a lock file, when the inventory runs, then every declared dependency appears with both versions and the counts match the files.
  • Given a repository with a manifest and no lock file, when the inventory runs, then it reports the absence as a single finding rather than one finding per dependency.

FR-TRI-02 — Suppressions expire

Priority: Should Requirement: An accepted-risk disposition shall carry an expiry date of no more than 180 days, and after that date the finding shall be reported as new again with the expired acceptance shown as context. Rationale: An acceptance with no expiry is a permanent decision made by a person who has since forgotten why. The expiry forces the decision to be renewed by somebody who is still around. Acceptance criteria:

  • Given an acceptance expiring next month, when the scan runs, then the finding is suppressed and the report shows the expiry date.
  • Given an acceptance recorded with an expiry more than 180 days out, when the triage file is validated, then validation fails naming the entry and the maximum.

FR-RPT-02 — The rotation runbook

Priority: Should Requirement: For every confirmed finding, the report shall link to a documented rotation procedure for that credential type, covering revocation, reissue, redeployment, and verification. Rationale: Finding the secret is the easy half. A key that is found and not rotated is still live, and a student who ships a scanner without a rotation procedure has built the half of the tool that feels good. Acceptance criteria:

  • Given a confirmed finding of a recognized credential type, when the report renders, then the four rotation steps for that type appear with the place each is performed named.
  • Given a confirmed finding of an unrecognized type, when the report renders, then the generic rotation procedure appears with an explicit instruction to identify the issuing system first.

Starter non-functional requirements

NFR-PERF-01 — Scan throughput

Priority: Must Requirement: A full-history scan shall complete in under 10 minutes on the reference machine named in docs/architecture.md, against a synthetic corpus of at least 20,000 commits produced by a committed generator script under script/ and listed in the scope file like any other target you own; and a full-history scan of the real in-scope estate shall complete inside the same bound. Measured by: three timed runs over the synthetic corpus and one timed run over the real in-scope repositories, median and worst case for each recorded in the measurements log each iteration. The 20,000 is a load figure for a generated corpus, not a claim about what you administer — FR-SCP-01 and the constraint above limit your real scope to what you own, so if your honest estate is four repositories and three thousand commits, that is the number, and it goes in the log beside the synthetic one. Do not pad the scope file with repositories you do not administer to reach a bigger figure; that is the one thing this brief forbids outright. Adapting a threshold to a real measurement window is exactly the skill this catalog is testing.

NFR-SEC-01 — No cleartext secret leaves the process

Priority: Must Requirement: No detected secret value shall appear in cleartext in any report file, log file, terminal output, temporary file, or pipeline output. Measured by: an automated test that scans a fixture repository containing twenty planted values, then searches every file the run wrote and both captured output streams for each planted value; the test fails on a single hit. This test runs in CI on every push.

NFR-QUAL-01 — Detection quality against a known corpus

Priority: Must Requirement: Against the committed fixture corpus of twenty planted secrets and at least two hundred decoy strings that resemble secrets but are not, the tool shall report at least eighteen of the planted values and at most three decoys. Measured by: an automated test asserting both counts, run in CI; the fixture corpus and its answer key are committed under tests/fixtures/ before the rule set is written.

NFR-REL-01 — Interruptible and idempotent

Priority: Must Requirement: A scan interrupted at any point and re-run shall produce the same finding set as an uninterrupted run, and shall never leave the triage file in a partially written state. Measured by: a test that kills the process at three points during a fixture scan, re-runs, and diffs the resulting finding set against the reference; plus a checksum assertion on the triage file after each kill.

NFR-ACC-01 — Severity readable without color

Priority: Should Requirement: Every severity in terminal output shall be identifiable from text alone, color shall be suppressed when the NO_COLOR convention is in effect or output is not a terminal, and any rendered HTML report shall meet a contrast ratio of at least 4.5:1 for body text. Measured by: capture piped output and confirm every severity is textually distinguishable with no escape sequences present; contrast checker over every token pair in the HTML renderer.

The genuinely hard part

False positives, and the fact that they are not a bug you can fix but a threshold you have to choose. Every rule that catches a real key also catches an example in a README, a test fixture, a base64-encoded image, and a hash that looks like a token. Tighten the rules and you miss real keys; loosen them and Amara wades through eighty findings, forty of which she has already dismissed twice, and stops running the tool — which is the actual failure mode, and it is a human one. NFR-QUAL-01 exists to turn that trade into a measurement instead of an argument, which means building the fixture corpus with its planted values and its decoys before you write the rules. Students build it afterward and it becomes a corpus that happens to pass.

The second hard problem is finding identity across a rewritten history. Your triage records are keyed to findings, and a finding is a value at a commit at a path — all three of which change when someone rebases, squashes, or rewrites history to remove the very secret you found. If your identifier is a commit hash, every triage decision evaporates the first time anyone cleans up a branch, and Amara re-triages forty findings and quits. Designing an identifier stable enough to survive that and specific enough not to collide is real engineering, it belongs in an ADR, and it is worth talking about in Week 8.

A stack that fits

A single-binary or single-script command-line tool in a language you already write well. You need fast file scanning, regular expressions, subprocess control, and hashing — every mainstream language has all four, so pick for fluency rather than fit. Shell out to git for history enumeration rather than binding a library, for the same reason as the previous brief: it is installed, it is correct, and it is one fewer unknown. Keep the triage file in a plain text format a human can read and edit in a code review, because reviewing a triage decision in a diff is most of the value.

Do not promise yourself a vulnerability advisory feed. Public advisory data sources exist, but their availability, licensing, and rate limits vary and change, and a capstone that discovers in Week 6 that its data source needs an account it cannot get is a capstone in trouble. If you want advisory matching, exercise the source for real in Week 2 — one live call, one saved response, one dated note about the terms in the assumptions section of docs/requirements.md — before it becomes a requirement. Everything in the Must list above works with no network at all, which is deliberate.

Recommendation, not verdict. Record the language, the git-subprocess decision, and above all the finding-identifier design as ADRs under docs/adr/ in Week 3.

What to cut first if Week 4 says you are behind

  1. Drop the dependency inventory; ship the secret half only. Saves about 10 hours and costs the project nothing conceptually — the triage, redaction, and reporting machinery is identical either way. Put it on the Won’t list with a reason.
  2. Drop the HTML report; ship terminal output plus a machine-readable file. Saves about 8 hours.
  3. Cut the rule set from twelve rules to five high-precision ones and say plainly in the README what the tool does not detect. Saves about 8 hours and usually improves NFR-QUAL-01.
  4. Cut suppression expiry and the orphaned-record handling, keeping plain triage. Saves about 6 hours. Cut this last — it is the part that makes the tool survive contact with a real repository.

If you are ahead

Add advisory matching over the dependency inventory, but only after you have verified the data source for real and recorded its terms with a date. About 15 hours including the mapping between how your ecosystem names a package and how the advisory source names it, which is where the time actually goes and which is a genuinely interesting problem to present. The alternative, if the data source does not work out, is a second scanning target from the same scope file — your own host’s configuration — using the identical triage and reporting path.


Brief 8 — Access Check (an accessibility audit-and-remediate workbench)

One line: A tool that audits a small set of web pages you are responsible for, records the manual checks automation cannot do, and tracks every barrier from discovery through fix to re-test.

The person who has this problem: Ana Villareal, the volunteer web coordinator for a forty-year-old community choir, who maintains a twenty-page site in her evenings and has just learned that Ruth — a board member of eleven years who uses a screen reader — has not been able to register for a concert since the site was redesigned.

Why it fits 160 hours: The page set is small and fixed, you control it, and the hardest input problem — getting the pages — is solved on day one. Weeks 1–2 are charter and requirements (40 h). Week 3’s skeleton audits one page with one automated rule and records one finding, in CI, at about 8 h of that week’s 20. The 40 construction hours of Weeks 5–6 split roughly: page collection and the audit harness 8; the finding model and store 6; the manual checklist capture 6; remediation tracking and the re-test difference 8; the report 5; preparing and running one real-user session 4; and 3 for integration and the defects that always appear. Weeks 7 and 8 are documentation, deployment, handoff, and delivery.

The problem

Ana is not indifferent and she is not unskilled. She is one person with a theme she did not write, and the redesign that broke Ruth’s registration also made the site look considerably better to everyone who can see it. Nobody told Ana anything was wrong for four months. When Ruth finally mentioned it, what Ana got was “the registration page doesn’t work with my screen reader,” which is true and is not actionable — she does not know whether that means the form has unlabeled inputs, or the error messages are announced as nothing, or the submit control is a styled div with no role, or all three.

So Ana did what most people do: she ran a free automated checker, got fifty-one issues, fixed the eleven she understood, and had no way to tell whether Ruth could now register. Four of the eleven were cosmetic. The one that was actually blocking Ruth — an error summary that appeared visually but was never announced — was not on the list at all, because no automated rule can tell whether a screen reader user learns that their submission failed.

That gap is the whole project. Automated checks are genuinely useful and genuinely partial: they find the machine-checkable subset of the barriers, and the tools’ own documentation says as much. A workbench that reports only the automated findings will hand Ana a number that goes to zero while Ruth still cannot register, and a number that goes to zero while the user is still blocked is worse than no number. What Ana needs is a tool that treats the manual pass and one real session with Ruth as first-class recorded data, refuses to call a page audited without them, and then tracks each barrier to a fix and a re-test.

Minimum viable scope (the Must list)

  • A declared page set and an automated run — audit only pages listed in a scope file, recording each finding against a named success criterion with the element and the evidence. (~8 h)
  • A structured finding record — page, element selector, criterion, evidence, and user impact, so that every finding can be sorted, assigned, and re-tested rather than read once in a wall of text. (~6 h)
  • A manual checklist captured as data — keyboard order, focus visibility, alt-text meaningfulness, error identification, and heading structure, recorded per page with a verdict and a note. (~6 h)
  • Remediation tracking with a re-test — each finding has a status and an owner, and a re-run reports fixed, still open, and newly introduced separately. (~8 h)

Four. Crawling the site to discover pages, applying fixes to the site itself, a hosted dashboard, auditing more than one site, and a second validation session measuring the change after remediation are not in the release.

The alternative shape for this brief is an assistive interface for one named person’s real need — a simplified, high-contrast, keyboard-and-switch-operable interface over a tool that person already has to use. It is a fine capstone and it is scoped the same way: one named person, three tasks, measured before and after. Its extra risk is that its acceptance criteria depend entirely on that person’s availability, so if you take it, get their calendar commitment in writing in Week 1 and name the risk in docs/risk-register.md.

Starter functional requirements

Ten drafted requirements. Ana and Ruth are sketches; every one of them needs a Source: line naming the real site owner and the real assistive-technology user you talked to before it earns a place in docs/requirements.md.

FR-SCP-01 — Audit only the declared page set

Priority: Must Requirement: The tool shall read a committed scope file listing every page to be audited, together with the host you are responsible for, and shall refuse to request any URL outside that list. Rationale: It keeps the audit bounded and repeatable, and it keeps you off hosts that are not yours. It also makes the run deterministic, which NFR-REL-01 depends on. Acceptance criteria:

  • Given a scope file listing twenty pages, when the audit runs, then exactly twenty pages are requested and the report names the scope file.
  • Given a page that redirects to a host not in the scope file, when the audit runs, then that page is recorded as unaudited with the redirect target as the reason, and no request is made to the other host.

FR-AUT-01 — Run the automated checks

Priority: Must Requirement: The tool shall run a documented set of automated checks against every in-scope page and record each violation as a finding, and the run shall complete and report even when individual pages fail to load. Rationale: This is the cheap, repeatable half of the audit. Its value comes entirely from being run again after every fix, which means it has to be one command. Acceptance criteria:

  • Given a page with an image lacking a text alternative and a control with insufficient contrast, when the audit runs, then two findings are recorded, each naming its success criterion.
  • Given one page that times out, when the run completes, then the other pages are reported normally, the timed-out page is recorded as unaudited with the cause, and the process exits non-zero.

FR-AUT-02 — The finding record

Priority: Must Requirement: Every finding, automated or manual, shall carry a stable identifier, the page, an element selector where applicable, the success criterion it relates to, the evidence observed, and a user-impact rating drawn from a documented scale. Rationale: Ana has an evening a week. Sorting by user impact is the difference between fixing the thing blocking Ruth and fixing four cosmetic contrast warnings that were easier. Acceptance criteria:

  • Given any finding, when the store is inspected, then all six fields are present, and the selector resolves to exactly one element on the page as captured.
  • Given a finding that applies to the page as a whole rather than an element — a missing page language, say — when it is recorded, then the selector field is explicitly empty rather than absent, and the report renders it as page-level.

FR-MAN-01 — Capture the manual pass

Priority: Must Requirement: An auditor shall be able to record, per page, a verdict and a note for every item on the committed manual checklist, covering at minimum keyboard operability and order, focus visibility, meaningfulness of text alternatives, programmatic labels on every input, error identification in text, and heading structure. Rationale: These are the checks that catch what actually blocked Ruth, and none of them can be automated. Storing them as data rather than prose is what lets them be re-tested and reported alongside the automated findings. Acceptance criteria:

  • Given an auditor completing the checklist for the registration page, when the entries are saved, then each item has a verdict and the failing items have generated findings with the same shape as automated ones.
  • Given an auditor who saves with three items unanswered, when the entry is saved, then the page’s manual pass is recorded as incomplete and the three unanswered items are named.

FR-FIX-01 — Track a finding to a fix

Priority: Must Requirement: Each finding shall carry a status of open, in progress, fixed, or accepted-with-reason, an owner, and the date of the last status change, and the history of status changes shall be retained. Rationale: Ana’s real workflow is a list she works through over six weeks. Without persistent status, every audit run resets her progress to zero. Acceptance criteria:

  • Given a finding moved to fixed, when the store is inspected, then the status, owner, and date are recorded and the previous status remains in the history.
  • Given a finding marked accepted-with-reason and no reason supplied, when the change is saved, then it is rejected and the required field is named.

FR-RPT-01 — The report Ana can act on

Priority: Must Requirement: The tool shall produce a report ordered by user impact and then by page, showing for each open finding the page, what a user experiences, the exact remediation, and the success criterion, in language a non-developer can act on. Rationale: The audience for this report is a volunteer with an evening a week, not an engineer. A report she cannot act on alone is a report that produces no fixes. Acceptance criteria:

  • Given a report containing findings at three impact levels, when Ana opens it, then the highest-impact findings appear first, each with a remediation sentence that names the file or component to change.
  • Given a page recorded as unaudited, when the report renders, then it appears in its own section with the reason, and is not reported as having zero findings.

FR-MAN-02 — A page is not audited until the manual pass is complete

Priority: Should Requirement: The tool shall report a page’s audit status as complete only when both the automated run and every manual checklist item have recorded verdicts, and shall label pages with automated results alone as partially audited. Rationale: This is the requirement that stops the tool from lying. A green automated result on a page nobody has driven from a keyboard means only that the machine-checkable subset passed. Acceptance criteria:

  • Given a page with a clean automated run and a complete manual pass, when the report renders, then the page is labelled fully audited with both dates.
  • Given a page with a clean automated run and no manual entries, when the report renders, then the page is labelled partially audited and the missing checklist items are named.

FR-FIX-02 — The re-test difference

Priority: Should Requirement: After a subsequent run, the tool shall report fixed findings, still-open findings, and newly introduced findings as three separate lists, matching findings across runs by their stable identifier. Rationale: Newly introduced is the list nobody builds and everybody needs. A redesign that fixes six barriers and creates two is not obviously an improvement, and only the difference shows it. Acceptance criteria:

  • Given a previous run with five findings and a current run in which two no longer occur and one is new, when the difference renders, then the three lists contain two, three, and one finding respectively.
  • Given a page whose markup changed such that a selector no longer resolves, when the difference is computed, then the finding is reported as needing re-verification rather than silently counted as fixed.

FR-VAL-01 — One real session, recorded

Priority: Should Requirement: The tool shall store, per validation session, the participant’s assistive technology, three attempted tasks, whether each was completed unaided, the time taken, and the point of failure, and shall link any barrier observed to a finding. Rationale: This is the only evidence in the whole project that the site works for the person it is supposed to work for. Everything else is a proxy. Acceptance criteria:

  • Given a recorded session in which one task failed, when the report renders, then the failure is shown with its linked finding at the highest impact level regardless of what any automated rule said.
  • Given a session where a task was completed only with assistance, when it is recorded, then it is stored as not completed unaided, with the assistance described.

FR-RPT-02 — The report meets the bar it sets

Priority: Should Requirement: The tool’s own report, in whatever form it is rendered, shall pass the tool’s own automated check set with zero findings and shall be fully operable by keyboard. Rationale: An inaccessible accessibility report is the single most embarrassing outcome available to this project, and a grader will check it in the first thirty seconds. Acceptance criteria:

  • Given a rendered report, when the tool is run against it, then zero findings are produced.
  • Given a report rendered with a finding table, when navigated by keyboard alone, then every control including any expand and sort controls is reachable and operable with a visible focus indicator.

Starter non-functional requirements

NFR-PERF-01 — Audit duration

Priority: Should Requirement: A full automated audit of twenty pages shall complete in under 3 minutes on the reference machine named in docs/architecture.md. Measured by: three timed runs against the fixture page set, median and worst case recorded in the measurements log each iteration.

NFR-REL-01 — Determinism

Priority: Must Requirement: Three consecutive audits of an unchanged page set shall produce identical finding sets, and no more than one finding per twenty pages may vary between runs. Measured by: a CI step that runs the audit three times over the committed fixture pages and diffs the finding identifiers; the step fails above the threshold. Flaky findings are logged in the defect log with the rule that produced them.

NFR-ACC-01 — The tool passes its own bar

Priority: Must Requirement: Every interface the tool presents — the report and any capture screen — shall be operable by keyboard alone with a visible focus indicator, shall meet a contrast ratio of at least 4.5:1 for body text and 3:1 for large text, shall convey no information by color alone, and shall have a programmatically associated label on every input. Measured by: unplug the mouse and complete the three core tasks, recording every place you got stuck; contrast checker over every text-and-background token pair; set the display to grayscale and repeat the three tasks; one screen-reader pass over the report. Results in the test results document.

NFR-PRIV-01 — Store the minimum

Priority: Must Requirement: The tool shall store, per finding, only the selector and a snippet of at most 200 characters; it shall never store full page bodies, and no personal data from any participant shall be committed to the repository beyond an initial and the assistive technology used. Measured by: a test asserting the snippet length bound on every stored finding, plus a review of the committed fixtures and session records before each milestone commit, recorded in the requirements change log.

NFR-SEC-01 — Requests stay in scope

Priority: Must Requirement: The tool shall issue network requests only to hosts named in the scope file, shall send no credentials, and shall not follow redirects to out-of-scope hosts. Measured by: an automated test running the audit against a fixture page that redirects and links off-host, asserting from the recorded request log that no out-of-scope host was contacted.

The genuinely hard part

Automation finds the machine-checkable subset, and the barrier that actually blocked Ruth is usually outside it. That is not a limitation you engineer around; it is the truth your tool has to be honest about, in its report, in its scoring, and in your Week 8 presentation. The design consequence is that the manual checklist and the validation session cannot be a documentation appendix — they have to be data flowing through the same finding pipeline as the automated results, with the same identifiers and the same re-test behavior. Building the automated half first and grafting the manual half on in Week 6 is the failure path, and it is the one nearly everybody takes, because the automated half is the fun half.

The second schedule risk is a calendar risk and it is severe: one real session with one real user, and that user is a person with their own life. Recruit in Week 1, confirm in Week 3, and schedule the session for Week 5 with a fallback date in Week 6 — because if the session slips to Week 7 you will have no time to fix anything you learn, and a finding you did not act on is worth a fraction of one you did. Name it in docs/risk-register.md with a trigger and a response. Treat the participant’s time as the scarcest resource in your project, prepare the three tasks in advance, and never ask someone to demonstrate their own difficulty for your benefit without being clear about what you are asking and what they get out of it.

One accuracy note that will cost you if you skip it: cite success criteria by number and check the numbers against the specification itself rather than a blog post or an assistant. Which version of the guidelines is current, and which criteria are at which conformance level, are facts that have changed and will change again. Write the version and the date you checked into the assumptions section of docs/requirements.md, and do not quote a statistic about what proportion of barriers automated checking catches unless you can name the study it came from.

A stack that fits

A headless browser driver plus an open-source accessibility rule engine, wired together by a script in whatever language you already know — Node and Python both have mature bindings for this kind of automation. There are several open-source rule engines; pick one, read its license and its rule documentation, and record both with the date in your ADR, because the rules a given engine implements are the rules your tool checks and you will be asked which those are. Store findings in SQLite or in committed structured text; both are fine, and text has the advantage that Ana can see a diff. Render the report as static HTML from the stored findings, which lets FR-RPT-02 be checked by running your own tool against your own output — a satisfying and genuinely useful test.

Novelty load is the thing to watch. Browser automation, a rule engine, a storage layer, and a report renderer is four moving parts, and browser automation in continuous integration is reliably more annoying than anyone expects — headless browsers need dependencies your runner may not have. Stand that up in Week 3 as part of the walking skeleton, not in Week 6. If it fights you for more than a day, fall back to auditing saved copies of the pages rather than live URLs, and record it as a scoping decision.

Recommendation, not verdict. Adopt or deviate in Week 3 with docs/adr/0001-choose-the-stack.md, and record the rule-engine choice and its license in its own ADR.

What to cut first if Week 4 says you are behind

  1. Cut the page set from twenty to nine — the pages that carry the three tasks that matter. Saves about 8 hours across audit runs, manual passes, and remediation. Nobody will fault a narrower, honestly stated scope.
  2. Drop the capture interface; record manual checklist entries in a committed structured file edited by hand. Saves about 10 hours and loses nothing an auditor cares about, because the auditor is you.
  3. Drop the re-test difference; report only the current state. Saves about 8 hours. Say clearly in the README that run-over-run comparison is out of scope, and put it on the Won’t list.
  4. Reduce the automated check set to the rules that map to the criteria on your manual checklist. Saves about 6 hours and makes the report easier to defend. Never cut the manual pass or the validation session — they are the only parts of this project that produce ground truth.

If you are ahead

Do the before-and-after properly: run the same three tasks with the same participant after remediation, and report the change in unaided completion and in time to complete, with an honest account of what a single participant can and cannot tell you. About 12 hours including scheduling, and it turns your Week 8 presentation from “I built an audit tool” into “Ruth can register, and here is the evidence” — which is a different talk entirely, and the one worth giving.


Brief 9 — Shift Deck (volunteer scheduling for one named nonprofit)

One line: A phone-friendly board where a small nonprofit posts the volunteer shifts it needs covered, volunteers claim and cancel them, and one coordinator can see at a glance which shifts are still short.

The person who has this problem: The volunteer coordinator at a small nonprofit — one part-time staffer, often twenty hours a week, who owns the schedule for forty to a hundred and fifty volunteers and currently runs it out of a group text, a shared spreadsheet, and her own memory. Not “volunteers.” Not “the organization.” One person, with a title, whose Sunday night is ruined by this.

Why it fits 160 hours: Four capabilities, all of them ordinary web work, none of them requiring a technology you have never touched. The 40 hours of Week 5–6 construction split across the four Musts: publishing shifts 7, claim-and-cancel 8, the coverage view 6, check-in and the hours number 7 — 28 hours of Must construction against the thirty-hour budget, with the remaining 12 held back for tests, integration, and the defect log. Week 3’s twenty hours pay for the design decisions and the walking skeleton — schema, one route end to end, tests running in CI, a real deploy — which is exactly why the Week 5–6 hours can be feature hours instead of setup hours. The unbudgeted risk is not code. It is the human being at the nonprofit, and the “genuinely hard part” section says so plainly.

The problem

Every small nonprofit that runs on volunteers runs on the same broken system. The coordinator posts next week’s needs in a group chat or a mass email. Six people reply “I can do Thursday.” Two of them mean different Thursdays. One replies to the coordinator directly instead of the group, so nobody else knows that slot is taken. By Wednesday the coordinator is texting individuals one at a time to find out whether the Saturday morning sort is actually covered, and she is doing it from her phone in a parking lot, because that is when she has a minute.

What makes it worse is that the information does not accumulate. At the end of the quarter, a board member asks how many volunteer hours the organization logged — a number that goes in grant applications and, for many organizations, has real money attached to it — and the honest answer is a guess reconstructed from a sign-in clipboard that somebody photographed and then lost. The coordinator knows the guess is low. She has no way to make it better.

This is not a hard technical problem. That is the point of putting it in a catalog for an eight-week course. It is a real problem, with a real person attached, and building it well will teach you more about authorization, data modeling, and handing software to a non-technical owner than a flashier project would. It will also give you the one thing capstone panels reward above almost everything else: a user who is not you, who you can name, who will say in Week 8 whether the thing works.

Two rules before you adopt this brief. First, you must name the organization in Week 1 and have talked to a human being there by Friday, with a date you can write down. A civic project with an imaginary client is worse than a personal project, because it looks like evidence and is not. Second, get the scope in writing — a paragraph, in an email, that both of you agree to — and put it in docs/charter.md. When the coordinator asks for a donor database in Week 6, that paragraph is the only thing that saves you.

Minimum viable scope (the Must list)

  • Publish the week. The coordinator creates shifts — role, start and end, location, how many people are needed — and volunteers can see them. (~7 h)
  • Claim and cancel. A volunteer takes an open spot from a phone, and can give it back, with the coordinator told when a cancellation is late. (~8 h)
  • Coverage at a glance. One screen showing the next fourteen days, every shift, filled versus needed, sorted so the shortfalls are on top. (~6 h)
  • Check-in and the hours number. Day-of attendance recorded against the shift, producing a monthly total the coordinator can hand to the board without reconstructing it from memory. (~7 h)

Four. Not six. Shift swapping between volunteers, reminder messages, a volunteer directory, background-check tracking, and a donor module are all things this organization genuinely wants and none of them are in the release.

Starter functional requirements

Ten drafted requirements, drafted rather than inherited. Every one needs a Source: line naming a real human at the organization before it goes into docs/requirements.md, and the set is deliberately Must-heavy because it is exactly the minimum viable scope — when you add your own requirements in Week 2, re-run MoSCoW so Musts land near half the document.

FR-SHIFT-01 — Publish a shift

Priority: Must Requirement: The coordinator shall be able to publish a shift specifying role, start time, end time, location, and the number of volunteers needed, after which it appears on the public shift list. Rationale: The whole system’s input. Everything else reads from this record. Acceptance criteria:

  • Given a signed-in coordinator on the new-shift form, when they submit a complete shift, then it appears on the volunteer-facing list within one refresh, showing role, date, time, location, and spots remaining.
  • Given a submission whose end time is not after its start time, when the coordinator submits, then the system rejects it and states which field is wrong; no shift is created.

FR-SHIFT-02 — Recurring shift series

Priority: Should Requirement: The coordinator shall be able to publish a shift that repeats weekly until a stated end date, and shall be able to change or cancel a single occurrence without affecting the rest of the series. Rationale: Most nonprofit schedules are the same five shifts every week; without repetition the coordinator retypes them and stops using the tool by week three. Acceptance criteria:

  • Given a weekly series of eight occurrences, when the coordinator cancels the third occurrence, then that occurrence disappears from the list and the other seven are unchanged.
  • Given a weekly series where one occurrence has had its start time edited, when the coordinator edits the series’ default start time, then the edited occurrence keeps its own time and every unedited occurrence takes the new one.

FR-SIGN-01 — Claim an open spot

Priority: Must Requirement: An identified volunteer shall be able to claim one spot on a shift that has spots remaining, and shall never be able to claim a spot on a shift that is full. Rationale: Overbooking is the failure the group chat already has; a scheduler that reproduces it is worthless. Acceptance criteria:

  • Given a shift needing four volunteers with three claimed, when a volunteer claims it, then the shift shows four of four and no longer accepts claims.
  • Given a shift with exactly one spot remaining, when two volunteers submit a claim within the same second, then exactly one claim succeeds, the other is told the shift filled, and the shift shows no more than its stated capacity.

FR-SIGN-02 — Cancel a claim, with lateness recorded

Priority: Must Requirement: A volunteer shall be able to release a spot they hold at any time before the shift starts, and the system shall mark the release as late when it occurs within a coordinator-configured number of hours before the start time. Rationale: People cancel. The coordinator does not need to prevent it; she needs to know early enough to backfill, and she needs late cancellations to be visible rather than remembered. Acceptance criteria:

  • Given a volunteer holding a spot on a shift 48 hours away and a lateness threshold of 12 hours, when they release it, then the spot returns to the open pool and the release is not marked late.
  • Given the same volunteer holding a spot on the same shift, when they release it 3 hours before the start, then the spot returns to the pool and the release appears in the coordinator’s coverage view marked late with the volunteer’s name.

FR-COV-01 — Fourteen-day coverage view

Priority: Must Requirement: The coordinator shall be able to view every shift starting within the next fourteen days with its claimed count against its needed count, ordered with the largest shortfall first. Rationale: This is the screen that replaces the Sunday-night phone calls. It is the product. Acceptance criteria:

  • Given eleven shifts in the next fourteen days of which three are short-staffed, when the coordinator opens the coverage view, then the three short shifts appear above the fully staffed ones with their shortfall shown as a number.
  • Given no shifts in the next fourteen days, when the coordinator opens the view, then it states that no shifts are scheduled and offers the create-shift action, rather than showing an empty page.

FR-CHK-01 — Day-of check-in

Priority: Must Requirement: The coordinator shall be able to record, for each volunteer claimed on a shift, whether they attended, and to record an attendance for a volunteer who arrived without claiming a spot. Rationale: Claimed is not attended, and the walk-up volunteer is the one the clipboard always loses. Hours reporting is only as good as this record. Acceptance criteria:

  • Given a shift with four claimed volunteers, when the coordinator marks three present and one absent, then the shift’s attendance record shows three attended and one no-show, each with the volunteer’s name.
  • Given a volunteer who is not on the claim list, when the coordinator adds them to the shift’s attendance, then they are recorded as attended for that shift and counted in that month’s hours.

FR-CHK-02 — Monthly hours report

Priority: Should Requirement: The coordinator shall be able to export, for a chosen calendar month, one row per volunteer per attended shift containing the volunteer’s name, the shift date, the role, and the shift’s scheduled duration in hours, as a comma-separated file. Rationale: This is the number that goes in grant reports. It is also the fastest way to prove the tool paid for itself. Acceptance criteria:

  • Given a month containing 30 attendance records, when the coordinator exports it, then the file contains 30 data rows plus a header, and the summed hours equal the sum of the scheduled durations of those shifts.
  • Given a month with no attendance records, when the coordinator exports it, then a header-only file is produced and the interface says the month is empty.

FR-ACC-01 — Volunteer access without a password

Priority: Must Requirement: A volunteer shall be able to reach their own claim and cancel actions by following a single-use link sent to the address the coordinator has on file, without creating or remembering a password. Rationale: Volunteers will not maintain an account for four shifts a year. Every password you require costs you real users, and every password you store is a liability you are not staffed to carry. Acceptance criteria:

  • Given a volunteer on the coordinator’s roster, when they request access, then a link is sent to their address on file that grants them their own view and no other volunteer’s.
  • Given an access link that has already been used or has passed its expiry window, when it is followed, then access is refused and a fresh link can be requested; no session is created.

FR-ACC-02 — Contact details are coordinator-only

Priority: Must Requirement: Volunteer email addresses and phone numbers shall be visible only to accounts holding the coordinator role; all other views shall identify a volunteer by display name only. Rationale: The organization is handing you a list of its supporters’ contact details. Leaking them is the one failure that ends the relationship and possibly the organization’s ability to recruit. Acceptance criteria:

  • Given a signed-in volunteer viewing a shift with three other volunteers claimed, when the page renders, then it shows three display names and no addresses or phone numbers, and the underlying response contains none either.
  • Given a volunteer session, when it requests a coordinator-only route directly, then the request is refused with an authorization error and the refusal is logged.

FR-SHIFT-03 — Text-message reminders

Priority: Won’t (this release) Requirement: Automated reminder messages to volunteers’ mobile phones before a claimed shift. Not built. Rationale: The coordinator will ask for this in Week 2 and it is genuinely the highest-value next feature. It is out because messaging to mobile numbers is a paid, registered, per-region dependency whose requirements and costs change — verify the current rules before you ever promise it — and because chasing it would cost the release its hours report. Revisit after handoff, with the organization paying for the channel.

Starter non-functional requirements

NFR-PERF-01 — Coverage view responsiveness

Priority: Must Requirement: The fourteen-day coverage view responds at a 95th percentile under 500 ms, against a database seeded with 18 months of history — at least 1,200 shifts and 6,000 attendance records. Measured by: 200 sequential requests from the script in script/, percentile reported in docs/test-results.md.

NFR-PRIV-01 — Contact-data minimization and retention

Priority: Must Requirement: The system stores exactly the personal fields the coordinator named as necessary — name, one contact address, optional phone — and nothing else, for production data at any time; attendance records older than 24 months are purged by a documented procedure. Measured by: a schema review recorded in docs/requirements.md §4 plus a demonstrated purge run in the runbook. Ask the organization whether it has its own data policy before you write yours; if it does, yours defers to it.

NFR-A11Y-01 — Volunteer screens meet WCAG 2.2 AA

Priority: Must Requirement: The two volunteer-facing screens — shift list and my shifts — return zero WCAG 2.2 Level AA violations from an automated checker, pass a complete keyboard-only run (claim and cancel a shift without touching a pointer), and meet the 4.5:1 minimum contrast ratio for normal text, on current stable versions of two browsers, one of them mobile. Measured by: automated report plus a recorded manual pass, both filed in docs/test-results.md. Check the current WCAG 2.2 text rather than my summary of it.

NFR-REL-01 — Capacity is never exceeded under concurrency

Priority: Must Requirement: Zero claims are ever accepted beyond a shift’s stated capacity, under 20 simultaneous claim requests against a shift with one spot remaining, repeated 25 times. Measured by: an automated concurrency test in tests/ that fails the build on a single overbooking.

NFR-OPS-01 — Restore drill

Priority: Should Requirement: The application is restored from a backup onto a clean environment in under 60 minutes, with attendance and shift counts matching the source exactly, by someone who is not you following the written runbook alone. Measured by: one dated drill recorded in docs/runbook.md before Week 8. The coordinator will outlive your involvement; her data has to.

The genuinely hard part

It is not the code. It is the recurrence model and the client relationship, in that order of surprise and reverse order of danger.

Recurring shifts look trivial and are not. The moment you allow “every Tuesday at 9” you have signed up for the exception problem: this Tuesday starts an hour later, that Tuesday is cancelled for a holiday, the whole series changes location from the eighth week on, and someone has already claimed a spot on an occurrence you are about to edit. Every calendar system in the world solves this the same way — store the rule, materialize occurrences, and let a materialized occurrence carry overrides that survive edits to the rule — and every student who improvises it instead spends eleven hours discovering that answer the expensive way. Decide it in Week 3, write it in an ADR, and if the estimate scares you, that is precisely why FR-SHIFT-02 is a Should.

The dangerous one is the coordinator. She is enthusiastic in Week 1 and unreachable in Week 6, not because she lost interest but because a grant deadline landed on her. Design for that from the start: get the scope agreed in writing in Week 1, book your Week 4 and Week 7 sessions on her calendar in Week 1 while she is enthusiastic, and make sure nothing on your critical path requires a response you cannot get. If the only person who can test check-in is unavailable the week check-in is due, you do not have a schedule risk, you have a schedule failure. Put her, by name, in docs/risk-register.md with a mitigation that does not depend on her.

A stack that fits

A conventional server-rendered web application in whatever language you are already fluent in, backed by a relational database, deployed to one managed host. Server-rendered, not a separate front-end application and API, because two deployables and a client-side router will cost you eight hours you need for recurrence and buy you nothing a volunteer will notice on a phone. Relational, because shifts, claims, and attendances are exactly the boring foreign-keyed shape relational databases were built for, and because FR-SIGN-01’s concurrency requirement is one line of transactional thinking in a real database and a research project without one. Passwordless email links for volunteers and one password-protected coordinator account, because that is the smallest authentication surface that satisfies FR-ACC-01 and FR-ACC-02.

Watch your novelty load. Count the technologies here that you have never shipped with. At two, you are fine. At three, cut a feature before Week 3 ends. At four, pick a different stack — this brief has no requirement that forces an unfamiliar one, and a capstone lost to a framework you were learning is the single most common way this course goes wrong.

This is a recommendation with reasoning attached, not a verdict. Deviate freely — but deviate in Week 3, in an architecture decision record in docs/adr/, with the alternatives you rejected and the reason. A defended deviation is worth more marks than a compliant default.

What to cut first if Week 4 says you are behind

  1. Recurring series (FR-SHIFT-02). Single-occurrence shifts only; the coordinator duplicates last week’s with one click. Saves 10–12 hours and removes the largest unknown in the project.
  2. The hours export (FR-CHK-02). Attendance is still recorded and visible on screen; the coordinator copies it. Saves 3–4 hours.
  3. Passwordless links (FR-ACC-01). Replace with one shared join code per organization plus a name field; contact details stay coordinator-only either way. Saves 6–8 hours, and costs you the ability to say a claim belongs to a verified person — write that trade-off down, because a grader will ask.
  4. Check-in (FR-CHK-01) collapses to a printable roster. The coordinator marks paper and types nothing. Saves 6–7 hours — the whole fourth Must, the hours number with it — and reduces the product to scheduling only. This is the cut that changes what the project is; make it in Week 4 with your instructor, not in Week 7 alone.

If you are ahead

Add role qualifications: volunteers hold zero or more qualifications, shifts require them, and the claim action refuses a volunteer who lacks one while telling them exactly which one and how to get it. It is a genuine authorization-modeling exercise, it is the feature every real volunteer coordinator asks for second, and it is honestly about 12 hours including the coordinator-facing screens to grant and revoke. Do not start it before your Week 6 verification work is finished.


Brief 10 — Interval (a retrieval-practice app for one real course)

One line: A drill app where a learner practices a deck of facts, grades each recall attempt, and gets each item back at a spacing interval chosen by a published algorithm — with a progress view that tells the truth about what is actually learned.

The person who has this problem: A student in one specific memorization-heavy course — the second-semester nursing student with three hundred drug names before a practical, the anatomy student with a hundred and twenty structures, the second-year Japanese student with the term’s kanji list. One course, one term, one deck, one deadline. Their secondary stakeholder is the instructor of that section, who currently cannot tell who is practicing and who is bluffing.

Why it fits 160 hours: The scheduler is a pure function over a review history — small, testable, and cheap to build once you have decided it in Week 3 instead of arguing with it in Week 6. The 40 construction hours split across the four Musts: deck import and editing 7, the practice loop 9, the scheduler and the backlog policy 8, the progress view 6 — 30 hours of Must construction against the thirty-hour budget, with the remaining 10 held back for tests, integration, and the defect log. Week 3 buys the schema, one end-to-end route, CI, and a deploy. What makes this fit is the discipline of building one deck type — text prompt, text answer — and refusing images, audio, typed-answer matching, and multiple choice until the Should list.

The problem

Retrieval practice with spacing is one of the better-replicated findings in learning research — the effect goes back to Ebbinghaus in 1885 and has been reproduced with wearying consistency ever since — and almost nobody studies that way, because doing it by hand is miserable. The paper-flashcard version requires you to physically manage boxes of cards on a schedule. The general-purpose flashcard sites do it well but treat every deck the same, tell your instructor nothing, and are not adaptable to how one particular course actually assesses.

So here is what actually happens in the memorization-heavy course. The student makes cards in week two, drills them hard for three days, feels good, stops, and then re-reads their notes for eleven weeks — because re-reading feels productive and testing yourself feels like failing. In week fourteen they cram. They pass or they do not, and either way they retain very little into the following term, which for a nursing student is a clinical problem and not just a grade problem.

The instructor sees the outcome and not the cause. She cannot tell the difference between a student who practiced steadily and got unlucky and one who never opened the material, and without that she cannot intervene in week five when intervening is still cheap. There is a version of this app that helps both of them, and there is a much more common version that quietly becomes a surveillance tool. Deciding which one you are building is part of this brief and it is graded.

One honesty rule, and it is not optional. You are building software that implements a scheduling algorithm. You are not demonstrating a learning gain. Claiming your app improves retention requires a controlled study you do not have eight weeks or an ethics approval to run, and writing that claim into your README is the kind of thing that gets found in a job interview. Specify what the software does; cite the literature for why the schedule is shaped that way; make no outcome claim you did not measure.

Minimum viable scope (the Must list)

  • Get the deck in. Import items from a plain comma-separated file and edit them, with bad rows reported by line number rather than silently dropped. (~7 h)
  • The practice loop. Show a due item, reveal the answer, take a grade on a fixed scale, move on — fast enough to be usable with a keyboard and no mouse. (~9 h)
  • The schedule. A documented, deterministic algorithm that computes the next due date from the item’s review history, plus a stated policy for what happens when a learner misses four days and the queue explodes. (~8 h)
  • Honest progress. A view showing what is due, what is scheduled far out, and what keeps failing — with each metric defined in writing before it is displayed. (~6 h)

Four. The instructor dashboard, typed-answer grading, images and audio, deck sharing, and streaks are not in the release, and the streak counter in particular is a behavioral-design decision you should have to defend before you build.

Starter functional requirements

Ten drafted requirements, for you to edit. Add your own Source: lines — for this brief the sources are one real learner in one real course and, if you involve them, one real instructor.

FR-DECK-01 — Import items from a file

Priority: Must Requirement: A signed-in learner shall be able to import practice items from a comma-separated file containing a prompt column and an answer column, and the system shall report every rejected row with its line number and the reason for rejection while importing all valid rows. Rationale: Nobody types three hundred items into a form. A silent partial import is worse than a failed one, because the learner discovers the gap during the practical. Acceptance criteria:

  • Given a 300-row file with 6 rows missing an answer, when the learner imports it, then 294 items are created and the report lists 6 rejections with their original line numbers and reasons.
  • Given a file whose header does not contain both required columns, when the learner imports it, then nothing is created and the system states which column is missing.

FR-DECK-02 — Edit and retire items without losing history

Priority: Should Requirement: A learner shall be able to correct an item’s prompt or answer, and to retire an item from future scheduling, without deleting that item’s review history. Rationale: Decks contain typos; typos are discovered mid-term. If correcting one resets its schedule or erases its history, learners stop correcting them and practice wrong facts. Acceptance criteria:

  • Given an item with 9 prior reviews, when the learner corrects its answer text, then the item’s due date and its 9 reviews are unchanged and the new text is what appears next time.
  • Given a retired item, when the learner starts a session, then it is never presented, and it still appears in the progress view marked retired with its history intact.

FR-SESS-01 — Run a practice session

Priority: Must Requirement: A learner shall be able to start a session that presents due items one at a time, reveal the answer, and record a grade on the fixed scale defined in the specification, entirely from the keyboard. Rationale: The loop is the product, and every extra interaction per item is multiplied by three hundred. Mouse-free operation is not polish here; it is the difference between a twelve-minute session and a twenty-five-minute one. Acceptance criteria:

  • Given a session with 20 due items, when the learner reveals and grades each one using only the keyboard, then all 20 grades are recorded and the session reports the count and elapsed time.
  • Given an item whose answer has not been revealed, when the learner attempts to grade it, then no grade is recorded and the interface states that the answer must be revealed first.

FR-SESS-02 — A session survives interruption

Priority: Must Requirement: When a session is interrupted, the system shall retain every grade already recorded and shall not present those items again in that day’s queue. Rationale: The learner practices on a phone between classes; the session is interrupted every time. Losing eleven grades to a closed tab is the defect that ends adoption. Acceptance criteria:

  • Given a session of 20 items with 11 graded, when the browser is closed and reopened, then the 11 grades are present in the history and the resumed queue contains the remaining 9.
  • Given the same interrupted session, when the learner grades the remaining 9, then the item count for that day is exactly 20 with no item counted twice.

FR-SCHED-01 — Deterministic next-due computation

Priority: Must Requirement: The system shall compute an item’s next due date solely from that item’s ordered review history and the current date, using the algorithm named and cited in the specification, such that identical histories and identical dates always produce identical due dates. Rationale: Determinism is what makes this testable. A scheduler you cannot test by simulation is a scheduler you can only test by waiting three weeks, and you do not have three weeks. Acceptance criteria:

  • Given the fixed set of 15 review histories committed to the repository, when the scheduler runs against each with a fixed clock, then every computed due date matches the expected value in the committed table.
  • Given any single history, when the scheduler is run over it 1,000 times, then all 1,000 results are identical.

FR-SCHED-02 — Backlog policy

Priority: Must Requirement: The system shall cap the number of items presented in a single day at a learner-configured limit, and when more items are due than the cap allows, shall present the most overdue items first. Rationale: Miss five days and a naive scheduler hands you 240 items and the learner quits. The cap is not a nicety; it is the difference between a tool that recovers from a bad week and one that punishes you for it. Acceptance criteria:

  • Given 240 due items and a daily cap of 60, when the learner starts a session, then exactly 60 items are queued and they are the 60 with the earliest due dates.
  • Given a daily cap of 60 and 47 due items, when the learner starts a session, then 47 items are queued and the session states that the deck is caught up.

FR-PROG-01 — Progress view with defined metrics

Priority: Must Requirement: A learner shall be able to view, for a deck, the count of items due today, the count scheduled beyond seven days, the count never yet answered correctly, and the count of items failed on three or more of their last five reviews, with each metric’s definition displayed alongside it. Rationale: Every number in a learning dashboard is a claim. Displaying the definition next to the number is what keeps this honest — and it is the section of the interface a panel will ask about. Acceptance criteria:

  • Given a deck with a known seeded history, when the learner opens the progress view, then each of the four counts equals the value computed independently by the test fixture.
  • Given any metric on the view, when the learner requests its definition, then the displayed definition matches the one written in docs/requirements.md §3.

FR-PROG-02 — Per-item history

Priority: Should Requirement: A learner shall be able to view, for any single item, every recorded review with its date, its grade, and the interval that was assigned as a result. Rationale: This is the debugging surface for the scheduler and the trust surface for the learner. When somebody says “this keeps coming back,” this screen answers whether that is a bug or the algorithm working. Acceptance criteria:

  • Given an item with 9 reviews, when the learner opens its history, then 9 rows appear in date order, each with grade and resulting interval.
  • Given an item with no reviews, when the learner opens its history, then the view states that the item has not been practiced and shows its scheduled first appearance.

FR-ACC-01 — Export and delete your own data

Priority: Should Requirement: A learner shall be able to export all of their own decks and review history as a machine-readable file, and to delete their account and all associated review history, with deletion taking effect within 24 hours. Rationale: You are storing a record of what a named student got wrong, repeatedly, with timestamps. Exit is the minimum decency, and it is also what lets you promise the instructor that the tool is not building a permanent file on anyone. Acceptance criteria:

  • Given a learner with 3 decks and 900 reviews, when they export, then the file contains all 3 decks and all 900 reviews and can be re-imported to reproduce the state.
  • Given a learner who confirms deletion, when 24 hours have passed, then no review record attributable to them remains in the primary store, verified by a query recorded in the runbook.

FR-SESS-03 — Free-text answer grading

Priority: Won’t (this release) Requirement: Grading a typed answer against the stored answer with tolerance for spelling and word order. Not built. Rationale: It is the feature learners ask for first and it is a fuzzy-matching project of its own — “acetylsalicylic” versus “acetyl salicylic” versus “asprin” is three different decisions, each of which needs its own evaluation set. Self-grading on a fixed scale is what the cited algorithms were designed around anyway. Revisit only with a written evaluation set and 15 spare hours.

Starter non-functional requirements

NFR-PERF-01 — The loop stays out of the way

Priority: Must Requirement: The elapsed time from a grade keystroke to the next prompt being fully rendered stays under 150 ms at the 95th percentile, on a deck of 5,000 items with 20,000 reviews of history, on the reference device named in the technical specification. Measured by: instrumented timings over a scripted 200-item session, reported in docs/test-results.md.

NFR-REL-01 — The scheduler is provably deterministic

Priority: Must Requirement: There are zero disagreements between the scheduler’s output and the committed expected-value table, and zero between repeated runs, across 10,000 simulated review sequences generated from a fixed seed and run against a fixed injected clock. Measured by: a property-style test in tests/ that fails the build on any disagreement. Building the injectable clock in Week 3 is what makes this possible at all; retrofit it in Week 6 and you will not.

NFR-PRIV-01 — Performance data is the learner’s

Priority: Must Requirement: An identifiable review record is readable by the learner it belongs to and nobody else, in all environments including your own development database; any instructor-facing view shows only aggregates over a cohort of at least five learners, with no per-item breakdown. Measured by: an authorization test per view plus a written data-flow paragraph in docs/requirements.md §4. If you deploy this for a real course with real students, ask the institution first — in the United States student education records fall under FERPA, and “it is a class project” is not an exemption you get to grant yourself.

NFR-A11Y-01 — The practice loop is operable without sight or a mouse

Priority: Must Requirement: The session screen returns zero automated WCAG 2.2 Level AA violations; the reveal and the grade are both announced by a screen reader; and no state — correct, incorrect, due, retired — is conveyed by color alone, on one screen reader and one desktop browser named in the test plan. Measured by: automated report plus a recorded manual pass in docs/test-results.md.

NFR-OPS-01 — A learner’s history survives you

Priority: Should Requirement: A restore from backup into a clean environment, performed by someone other than the author following docs/runbook.md alone, completes in under 60 minutes and yields review counts and due dates identical to the source. Measured by: one dated drill before Week 8.

The genuinely hard part

Time. Not the schedule — the clock. This is a system whose entire behavior is a function of dates, which means you cannot test it by using it. If your scheduler reads the current time from wherever it happens to be, your only verification is to wait until Thursday, and you will do that exactly twice before Week 6 eats you. The move is to build the scheduler as a pure function — review history plus a passed-in date goes in, a due date comes out — with no database access, no clock access, and no I/O of any kind inside it. Then a thousand simulated learners over a simulated year runs in under a second. Students who do this ship a correct scheduler in nine hours. Students who do not spend twenty and still are not sure.

The second hard part is choosing the algorithm and being honest about the choice. Implement something published and citable — the Leitner box system (Sebastian Leitner, 1970s) is the simplest defensible option and is genuinely adequate for a term-length deck; the SM-2 algorithm from SuperMemo, described publicly in the late 1980s and reimplemented many times since, is the classic interval-and-ease-factor approach; there are more recent open scheduling algorithms too. Read the actual published description, not a blog summary, cite it in an ADR, and write down what it does badly. What you must not do is tune an interval curve by feel and describe it as evidence-based. That is the one sentence in this project that a knowledgeable panelist will stop you on, and they will be right to.

A stack that fits

Any web stack you already know, a relational database, and one deliberate architectural rule: the scheduler lives in its own module in src/ with no imports from your database layer, your web framework, or your date-of-today helper. Everything else — sessions, imports, the progress view — is ordinary reads and writes.

Keep it a web application. A native mobile app doubles your deployment and testing surface for a loop that is four keystrokes long, and a phone browser handles this fine. If you want it on a home screen, a manifest is an hour of work and does not obligate you to the offline machinery in Brief 11 — but be honest in your ADR about which of those two you are doing, because “offline-capable” written casually into a charter is how a nine-hour project becomes a thirty-hour one.

Novelty load: at most two unfamiliar technologies. This brief’s difficulty is in modeling and testing, not in tooling, so spending your unfamiliarity budget on a new framework here is a poor trade. As always this is a recommendation to adopt or deviate from in Week 3, in an ADR in docs/adr/, with the alternatives and the reason — not a verdict.

What to cut first if Week 4 says you are behind

  1. Import (FR-DECK-01) becomes a paste box. One item per line, prompt and answer separated by a tab, no file handling, no rejection report. Saves 5–6 hours.
  2. Per-item history (FR-PROG-02) and two of the four progress metrics. Keep due-today and never-correct; drop the rest with the definitions written down as future work. Saves 5–6 hours.
  3. Export and delete (FR-ACC-01) becomes a documented manual procedure the operator runs on request, written into the runbook. Saves 6–7 hours and is honest as long as the README says so.
  4. Drop accounts entirely. One deck, one device, local storage, no sign-in. Saves 10–12 hours, kills NFR-PRIV-01 as a concern, and changes the project from a service into a tool. Make this call with your instructor in Week 4; it is a real project, but it is a different one.

If you are ahead

Build a scheduler comparison harness: implement a second algorithm behind a configuration flag, generate synthetic learners with parameterized forgetting behavior, and run both schedulers over a simulated term, reporting reviews required to reach a stated retention level. Roughly 14 hours, and it turns your Week 8 talk from “I built a flashcard app” into “here is what happened when I compared two published algorithms under stated assumptions.” Two warnings: your synthetic learner is a model, so state its assumptions loudly and claim nothing about real humans from it; and do not run the comparison on your classmates, because an experiment on real students needs approval you do not have.


Brief 11 — Round Trip (an offline-first field-round app)

One line: A phone-installable web app that lets a technician complete an inspection round in a basement with no signal, then syncs every record exactly once when the signal comes back — including when the sync is interrupted halfway and somebody else edited the same asset.

The person who has this problem: A facilities maintenance technician who walks a preventive-maintenance route — mechanical rooms, basements, roof plant, a rural pump station — where there is no cell coverage and no guest wireless, carrying a phone and a clipboard, and retypes the clipboard into a spreadsheet at the end of the day. Substitute your own domain freely: a well-water sampler, a stage-rigging inspector, a bike-share mechanic, a habitat surveyor. The shape does not change.

Why it fits 160 hours: Because the feature list is small and the difficulty is concentrated in one place, which is exactly what an eight-week project wants. The 40 construction hours split across the four Musts: offline capture and the local queue 9, the sync protocol and its server side 11, conflict detection and resolution 6, the supervisor view 4 — 30 hours of Must construction against the thirty-hour budget, with the remaining 10 held back for tests, integration, and the defect log. Week 3 buys the schema, the walking skeleton, CI, and — critically — the fault-injection harness, which is not optional here and is the reason Week 3 is not spare time. If you cut anything, cut features. Do not cut the harness. This is the tightest Must set in the catalog: it sits exactly on the thirty-hour line, with two of the four Musts carrying the hardest work in the book. Read the cut list below knowing that — its first two entries are Shoulds, so they buy back hours only if you adopted them in the first place, and the first cut that actually moves the thirty is the third one, which is a Must and therefore a conversation with your instructor rather than a decision you make alone on a Saturday.

The problem

The technician’s round is thirty stops. At each one there are eight to fifteen checks — a pressure reading, a belt condition, a leak yes-or-no, a photo when something looks wrong. She has a phone in her hand and no bars, because she is standing in a concrete room under a building. So the round happens on paper, and at 4:30 she sits at a desk and retypes ninety minutes of readings from her own handwriting.

Two things go wrong, reliably. Readings get transcribed wrong — a 2 becomes a 7, a stop gets skipped in the retyping and nobody notices for a month. And the retyping does not always happen. On a bad day it waits until tomorrow, and tomorrow it waits until Friday, and by Friday the sheet is a guess. Meanwhile the supervisor has no idea which of the four technicians finished their rounds until the data eventually shows up, which means a missed critical reading is discovered in the worst way.

Every off-the-shelf answer to this fails in the same place: it assumes connectivity. A form that requires a network is a form that requires the technician to walk back outside, and she will not. Which means the entire value of the project sits behind one requirement — this must work with the radio off — and everything hard about the project follows from that one requirement.

Minimum viable scope (the Must list)

  • Capture with the radio off. A full round completes offline, and the data survives the app being closed, the browser being killed, and the phone being restarted. (~9 h)
  • Sync exactly once. When the network returns, everything uploads; an interrupted upload resumes; nothing duplicates; nothing silently disappears; the technician can see exactly what is still pending. (~11 h)
  • Resolve conflicts on purpose. Detect that the server’s version has changed since the device’s copy, apply a written policy, and put a human in the loop when the policy says a human is required. (~6 h)
  • Supervisor visibility. A web view of completed rounds and, per device, when it last synced — because “did Maria’s round come in?” is the question the whole thing exists to answer. (~4 h)

Four. Route assignment, scheduling, work orders, asset history charts, and offline photo annotation are not in the release.

Starter functional requirements

Ten drafted requirements. Note that FR-SYNC-01 and FR-CONF-01 are where the project’s actual engineering lives; edit their acceptance criteria carefully rather than loosening them, because loosening them is how this brief turns into an ordinary form app with an offline sticker on it.

FR-CAP-01 — Complete a round with no network

Priority: Must Requirement: A technician shall be able to open the application, load an assigned round, and record a result for every check on every asset in that round while the device has no network connectivity. Rationale: This is the requirement the project exists for. If it is conditional on anything, there is no project. Acceptance criteria:

  • Given a device with all radios disabled and a round of 20 assets previously loaded, when the technician records results for all 20, then every result is stored on the device and the interface shows the round complete and pending sync.
  • Given a device with no network that has never loaded the application before, when the technician opens it, then they are told plainly that the round must be loaded while connected, rather than being shown an empty or broken screen.

FR-CAP-02 — Attach a photo within a size budget

Priority: Should Requirement: A technician shall be able to attach up to three photographs to any check result, and the system shall reduce each stored image to at most the byte budget stated in the specification before writing it to the device. Rationale: Photos are the highest-value evidence and the fastest way to blow your device storage. A phone camera image is several megabytes; forty of them will exceed what a browser will reliably keep for you. Acceptance criteria:

  • Given a 6 MB camera image, when the technician attaches it, then the stored image is at or under the stated budget and is still legible enough to read a nameplate at arm’s length.
  • Given a round that has reached the configured on-device image budget, when the technician attaches another photo, then the system states the budget is reached and names which rounds must be synced to free space, rather than failing silently.

FR-STORE-01 — Captured data survives restart

Priority: Must Requirement: Every captured result and queued upload shall remain intact and pending after the application is closed, the browser process is terminated, and the device is restarted, until it has been confirmed accepted by the server. Rationale: The technician’s phone dies mid-round approximately once a quarter. Losing a round to a battery is the defect that ends trust in the tool permanently. Acceptance criteria:

  • Given 14 unsynced results on a device, when the browser is force-quit and the device restarted, then all 14 are present and still marked pending.
  • Given a result that the server has confirmed accepted, when the device restarts, then the result is not re-queued and is not uploaded a second time.

FR-SYNC-01 — Each record uploads at most once

Priority: Must Requirement: The system shall ensure that repeating the upload of a record that the server has already stored creates no additional record and returns the same result as the original upload. Rationale: Retries are not an edge case; they are the normal operating mode of a system that syncs over flaky mobile networks. Duplicates in an inspection record are indistinguishable from a technician inspecting something twice, which is a lie in a compliance document. Acceptance criteria:

  • Given a record that has been successfully uploaded, when the same upload is repeated 5 times, then the server holds exactly one record and each response reports it as already accepted.
  • Given an upload that reaches the server and whose response never reaches the device, when the device retries, then the server holds exactly one record and the device marks it synced.

FR-SYNC-02 — Partial sync is safe and resumable

Priority: Must Requirement: When a sync of multiple records is interrupted, the system shall retain per-record status such that already-accepted records are not resent, unsent records remain queued, and the technician is shown how many of each remain. Rationale: All-or-nothing sync means one bad record blocks a whole round forever, and a sync with no per-record status means the technician cannot tell whether it is safe to close the app. Acceptance criteria:

  • Given a queue of 20 records where the connection drops after 12 are accepted, when the technician syncs again, then exactly 8 records are transmitted and the server holds 20.
  • Given a queue containing one record the server permanently rejects, when sync runs, then the other records are accepted, the rejected one is shown with the server’s stated reason, and it does not block subsequent syncs.

FR-SYNC-03 — Sync state is visible

Priority: Should Requirement: The technician shall be able to see, at any time and without a network, the count of records pending upload, the count failed with reasons, and the time of the last successful sync. Rationale: Trust in an offline tool is entirely a function of whether the user can tell what has left the device. Without this screen they will keep the paper backup, and if they keep the paper backup you have not replaced anything. Acceptance criteria:

  • Given 8 pending and 1 failed record, when the technician opens the sync screen offline, then it shows 8 pending, 1 failed with its reason, and the timestamp of the last successful sync.
  • Given a device that has never synced, when the screen is opened, then it states that no sync has occurred rather than showing a blank or epoch timestamp.

FR-CONF-01 — Conflicts are detected and resolved by a written policy

Priority: Must Requirement: When a device uploads a change to an asset record whose server-side version differs from the version the device last received, the system shall detect the divergence and apply the resolution policy stated in the specification, recording both the accepted and the rejected value in an audit record. Rationale: Two technicians on one asset, or a supervisor editing while a technician is offline, is not hypothetical — it is Tuesday. Undetected divergence means the last device to sync silently erases someone else’s work, and nobody finds out. Acceptance criteria:

  • Given a device holding version 3 of an asset and a server holding version 5, when the device uploads a change, then the divergence is detected, the policy is applied, and an audit record contains both values with their timestamps and authors.
  • Given any conflict resolved automatically, when a supervisor views that asset’s audit record, then the losing value, its author, and the rule that decided the outcome are all readable.

FR-CONF-02 — Human resolution when the policy defers

Priority: Should Requirement: When the resolution policy classifies a conflict as requiring human judgment, the system shall present both versions side by side to a supervisor with their authors and timestamps, and shall not apply either until a choice is made. Rationale: Some fields must not be resolved by a rule — a safety-critical reading is one of them. Automatic resolution of those is a decision to be silently wrong. Acceptance criteria:

  • Given a conflict on a field marked as requiring judgment, when it is detected, then the asset shows as awaiting resolution and neither value is applied.
  • Given a supervisor choosing one of the two versions, when they confirm, then that value is applied, the other is retained in the audit record, and the asset leaves the awaiting-resolution state.

FR-SUP-01 — Supervisor round view

Priority: Must Requirement: A supervisor shall be able to view every round for a chosen date with its completion status, and, per registered device, the time of that device’s last successful sync. Rationale: This is the question the organization actually has. It is also your only way to notice that a technician’s device has not synced in three days, which is the failure mode this whole architecture creates. Acceptance criteria:

  • Given four rounds on a date, two complete and two partial, when the supervisor opens that date, then all four appear with completion counts, and each is attributed to a technician.
  • Given a device whose last successful sync is more than 24 hours ago, when the supervisor opens the view, then that device is flagged with the elapsed time, by a means that does not rely on color alone.

FR-CAP-03 — Background sync while the app is closed

Priority: Won’t (this release) Requirement: Uploading queued records while the application is not open. Not built. Rationale: Support for background execution and push varies by browser and operating system and changes over time — verify what your user’s actual device does before you ever design around it. Depending on it would put your Must list on a capability you cannot guarantee. This release syncs while the app is open, and FR-SYNC-03 makes that state honest to the user. Revisit only with a verified device matrix.

Starter non-functional requirements

NFR-REL-01 — Zero loss, zero duplication under fault injection

Priority: Must Requirement: Zero records are lost and zero are duplicated across a 25-case fault matrix covering connection dropped mid-request, connection dropped before response, server 500, server timeout, app killed mid-sync, and device restarted mid-sync — each case run twice. Measured by: an automated harness in tests/ that reconciles device queue against server contents after every case and fails the build on any discrepancy.

NFR-PERF-01 — A round syncs on a bad connection

Priority: Must Requirement: A completed round of 20 assets including 10 photographs syncs in under 90 seconds on a simulated 1 Mbit/s uplink with 300 ms round-trip latency, using the throttling profile named in the test plan; and capture-to-stored-locally stays under 400 ms at the 95th percentile, because the technician is standing up holding a flashlight. Measured by: three timed runs recorded in docs/test-results.md.

NFR-A11Y-01 — Usable one-handed, in gloves, in bad light

Priority: Must Requirement: The capture screen returns zero automated WCAG 2.2 Level AA violations; every interactive target meets the WCAG 2.2 minimum target-size criterion at Level AA (24 by 24 CSS pixels — check the current specification text, and note the enhanced Level AAA criterion is larger at 44 by 44, which is the better target for a gloved thumb); no status is conveyed by color alone; and normal text meets a 4.5:1 contrast ratio — all on the physical device the technician actually carries, named in the test plan. Measured by: automated report plus one recorded field pass with the real user.

NFR-SEC-01 — Device data is minimized and expires

Priority: Must Requirement: Measured against the case of a lost or stolen phone: the device retains no personal data beyond the technician’s own identifier; synced records are purged from the device within 7 days of confirmed acceptance; all transport is over TLS; and session credentials expire within the window stated in the specification and are revocable by a supervisor. Measured by: an on-device inspection procedure written into docs/runbook.md and demonstrated once, plus an automated test for purge behavior. Photographs of a plant room can contain serial numbers, badge numbers, and people; say in writing what yours may contain and what you do about it.

NFR-OPS-01 — Storage pressure fails loudly

Priority: Should Requirement: Under a simulated write failure and a simulated eviction of the application’s stored data, the technician is told in words before starting a round, and no capture ever appears to succeed while failing to persist. Measured by: two scripted cases in the fault matrix. Browsers and operating systems may reclaim storage from web applications under pressure, and the specifics differ by platform and change over time — design as though eviction can happen, and verify current behavior on your target device rather than trusting any single account of it.

The genuinely hard part

Sync is the whole project, and the reason it is hard is that it is a distributed-systems problem wearing a form’s clothing. Every failure that matters happens between the device and the server: the request that arrived but whose response did not; the batch that was half accepted; two devices that were both right about the world at different times; a device clock that is four minutes fast, so “latest timestamp wins” quietly means “whoever’s phone is wrong wins.” None of these are visible from the interface, none of them show up in casual testing, and all of them show up in production in week one.

There is one design decision that removes most of the pain, and you should make it in Week 3: make your captured data append-only. A check result is an observation — this technician, this asset, this check, this value, at this time — not a mutable row. Two observations of the same thing are not a conflict; they are two observations, and the read model decides which one to show. That collapses FR-CONF-01 down to the genuinely mutable things — the asset’s own attributes, the round’s assignment — which is a much smaller surface, and it makes FR-SYNC-01 almost free, because inserting the same immutable observation twice is trivially detectable. Students who model observations as editable rows spend fifteen hours on conflict resolution and are still not confident. Write the ADR either way, but know that this is the fork in the road.

The second thing that will eat your schedule is that you cannot test this by hand. You need a harness that can script “accept 12 records, then drop the connection, then restart the app” and assert the outcome. Build it in Week 3, before the features. It feels like a detour and it is the single highest-return twelve hours in the project.

A stack that fits

An installable web application: a service worker so the application shell loads with no network, a structured client-side database for the queue and the captured records, and a small server exposing an upload endpoint plus the supervisor views. Web rather than native, because you get one codebase, an instant install path with no store review, and a deployment story you can actually finish in Week 7. The server can be anything you know; there is very little of it, and it should stay that way.

Be brutally honest about novelty load here, because this is the brief where it kills people. Service workers, a client-side database, a sync protocol, and conflict resolution may all be new to you — that is four, and four is a project you will not finish. Two, at most, should be unfamiliar. If service workers and the client-side store are both new, then your server, your language, and your test framework must all be old friends. If you are unwilling to make that trade, take a different brief; there is no shame in it and it is a better decision than a Week 7 rewrite.

One more warning specific to this brief: verify capabilities on the device your user actually carries, in Week 1. Installability, storage behavior, camera access, and background execution differ across browsers and operating systems and they change. Do not design around a capability you read about; design around one you personally saw work on that phone, and write the date you checked it into docs/requirements.md §4. As always this stack is a recommendation to adopt or deviate from in Week 3 with an ADR in docs/adr/, not a verdict.

What to cut first if Week 4 says you are behind

  1. Photographs (FR-CAP-02). Text and numeric results only. Saves 10–12 hours — image capture, resizing, storage budgeting, and multipart upload are four problems, not one — and it is by far the highest-value cut available to you.
  2. Human conflict resolution (FR-CONF-02). Keep detection and the audit record; resolve everything by the stated automatic policy and show a “your change was superseded” banner. Saves 7–8 hours and keeps the interesting part.
  3. The supervisor view (FR-SUP-01) becomes an export. A single authenticated endpoint that emits the day’s rounds as a file. Saves 3–4 hours — and because FR-SUP-01 is a Must, this one is a scope change you take to your instructor rather than decide alone.
  4. One device per technician, enforced. Drop multi-device reconciliation entirely and say so in the constraints. Saves 3–4 hours and narrows what FR-CONF-01 has to handle.

If you are ahead

Add a server change feed and two-way sync: the device pulls asset and round changes made since its last sync, so a supervisor’s edit reaches the technician’s phone without a reinstall. Roughly 16 hours done properly, and it is a real step up in difficulty because you now need a stable ordering of server changes and a device-side cursor that survives interruption exactly as carefully as your upload queue does. Do not start it before the fault matrix in NFR-REL-01 is green.


Brief 12 — Restore First (a backup-and-verify tool for a small self-hosted service)

One line: A command-line tool that backs up one small service on a schedule, then proves each backup by actually restoring it into a scratch location and checking it, so the operator can answer “when was the last backup I know is good?” with a date instead of a hope.

The person who has this problem: The one volunteer operator who runs a small self-hosted service — a club’s forum, a lab’s instrument database, a church’s scheduling app, a professor’s research web app — on a single machine, alone, unpaid, whose entire disaster-recovery plan is a scheduled job that writes files to a second disk and has never once been restored.

Why it fits 160 hours: A command-line tool has no interface to design, no browser matrix, no accessibility audit of a layout, and no deployment story beyond “it runs where the service runs.” All of that budget goes into correctness. The 40 construction hours split across the four Musts: the backup path and its manifest 8, verification by rehearsal restore 10, restore with dry-run and refusal-to-clobber 7, the operator’s status and report 5 — 30 hours of Must construction against the thirty-hour budget, with the remaining 10 held back for the fault-injection matrix, integration, and the defect log. Week 3 buys the design decisions, the walking skeleton — one source, one destination, one round trip through backup and restore in CI — and the disposable test target that everything else depends on.

Scope note, stated once and meant: everything in this brief runs against systems you own or have written permission to operate. You are backing up your own service, restoring to your own scratch host, and breaking your own containers on purpose. This is defensive work on your own infrastructure, full stop. If your project touches a machine somebody else owns, get it in writing before Week 2 and put the permission in the repository.

The problem

The operator set this up two years ago. There is a job that runs at 3 a.m. and copies things to /backups, and it has run every night since, and the disk light blinks, and everybody feels fine. Nothing about that arrangement is a backup. It is a copy process with an unknown success rate, and there are at least four ways it is probably already broken: the copy runs while the database is being written to, so what is on the disk is a torn file that will not open; the destination filled up in March and the job has been failing silently ever since because nobody reads the output of a cron job; the retention script deleted the wrong generation once and nobody noticed; and nobody has ever tried a restore, so nobody knows that restoring requires a version of a tool that is no longer installed anywhere.

The operator finds all of this out on the day it matters, which is the day the disk fails, and the thing they lose is not their weekend. It is a decade of a club’s records, or a research group’s data, or a hundred families’ registrations.

What separates a backup from a copy is exactly one property: somebody has restored it and checked the result. That is not a feature you bolt on; it is the entire architecture. Which is why this brief is not called “a backup tool.” Build the verification first, build the restore second, and the backup last, and you will end up with something a real operator can actually rely on — and with a Week 8 demonstration that is genuinely gripping, because you get to break the system live and show it refusing to lie.

Minimum viable scope (the Must list)

  • Take a consistent backup. A defined source set — one database plus one directory tree — captured in a way that is internally consistent, with a manifest recording what was captured, when, how big, and what it hashes to. (~8 h)
  • Prove it by restoring it. Every backup is automatically restored into a disposable scratch target and checked, and the result — pass or fail, with the reason — is recorded against that backup. (~10 h)
  • Restore for real, safely. One command restores a chosen backup to a chosen target, with a dry run that shows exactly what would happen and a refusal to overwrite anything without an explicit flag. (~7 h)
  • Tell the operator the truth. A status the operator actually sees: a non-zero exit on failure, a one-command answer to “when was the last verified good backup,” and an alert when runs fail repeatedly. (~5 h)

Four. Encryption, deduplication, incremental backups, a web dashboard, and multi-host orchestration are not in the release.

Starter functional requirements

Ten drafted requirements, for editing. FR-VER-01 is the requirement that makes this project worth doing; if you weaken it into a checksum comparison, you have built a copy tool with extra steps.

FR-BKUP-01 — Produce a consistent backup with a manifest

Priority: Must Requirement: The tool shall produce, from the source set named in its configuration, a single dated backup artifact accompanied by a manifest recording the source set, the start and end time, the total byte count, and a cryptographic digest of the artifact. Rationale: Without a manifest you cannot verify, prune, or report on anything. The digest is what makes “this file is intact” a checkable claim rather than an assumption. Acceptance criteria:

  • Given a configured source of one database and one directory, when a backup runs, then one artifact and one manifest are produced, and the manifest’s digest matches a digest independently computed over the artifact.
  • Given a running database receiving writes throughout the backup, when the backup is later restored, then the restored database opens without error and passes the configured structural check. (Copying a database’s files while it runs does not satisfy this. Use the database’s own dump or online-backup mechanism, or a filesystem snapshot — and name which, in an ADR.)

FR-BKUP-02 — A partial backup is never marked good

Priority: Must Requirement: When a backup run fails at any point, the tool shall leave no artifact recorded as complete, shall retain or remove the partial artifact according to the configured policy, and shall exit with a non-zero status. Rationale: The failure that actually destroys data is a half-written archive that the retention rule counts as a generation and the operator counts as safety. Acceptance criteria:

  • Given a destination that runs out of space at 60% of the way through, when the backup runs, then no manifest is written marking the run complete, the exit status is non-zero, and the failure names the destination and the reason.
  • Given a backup process terminated mid-write, when the tool next runs, then the incomplete artifact is identified as incomplete and is never selected for restore or counted for retention.

FR-VER-01 — Verify by rehearsal restore

Priority: Must Requirement: After each successful backup, the tool shall verify it by confirming the artifact’s digest, restoring it into a disposable scratch target, and executing the structural checks named in the configuration, recording a pass or fail with the reason against that backup. Rationale: This is the requirement that turns a copy into a backup. Everything else in this tool is plumbing around it. Acceptance criteria:

  • Given a good backup and configured structural checks (for example: the database opens, a named table exists, its row count is within the configured tolerance of the source, and a named file is present with matching size), when verification runs, then every check passes and the backup is recorded verified with the timestamp and the check results.
  • Given a backup artifact whose bytes have been altered after creation, when verification runs, then the digest check fails, the backup is recorded as failed with the reason, and the exit status is non-zero.

FR-VER-02 — Answer “when was the last good one?”

Priority: Should Requirement: The tool shall report, on demand, the identifier and completion time of the most recent backup that passed verification, and the age of that backup in hours. Rationale: This is the operator’s only question, and it must be answerable in one command at 2 a.m. without reading a log. Acceptance criteria:

  • Given five backups of which the newest two failed verification, when the operator requests the status, then the third-newest is reported as the last verified good, with its age in hours.
  • Given no backup has ever passed verification, when the operator requests the status, then the tool says so explicitly and exits non-zero, rather than reporting nothing.

FR-REST-01 — Restore with a dry run

Priority: Must Requirement: The operator shall be able to restore a chosen backup to a chosen target, and shall be able to request a dry run that reports every action that would be taken — files written, database objects replaced, bytes moved — without changing anything. Rationale: The restore command is run once a year, under stress, by someone who has forgotten the flags. The dry run is how they find out what it will do before it does it. Acceptance criteria:

  • Given a verified backup and an empty target, when the operator runs a dry run, then the tool lists the actions and the byte count, changes nothing at the target, and exits zero.
  • Given the same backup and target, when the operator runs the real restore, then the actions taken match the dry run’s list, and the restored data passes the same structural checks used in verification.

FR-REST-02 — Refuse to clobber without being told

Priority: Must Requirement: When the restore target already contains data, the tool shall refuse to proceed and shall exit non-zero, unless the operator has supplied the explicit overwrite flag named in the documentation. Rationale: The second-worst outcome in this project is failing to restore. The worst is restoring over the live system by accident, at which point your recovery tool has become the incident. Acceptance criteria:

  • Given a non-empty target, when restore is invoked without the overwrite flag, then nothing is written, the exit status is non-zero, and the message names what was found and which flag would proceed.
  • Given a non-empty target and the explicit overwrite flag, when restore is invoked, then the restore proceeds and the run record states that an overwrite was performed, on what, and by whom.

FR-RET-01 — Prune by policy, never the last good backup

Priority: Should Requirement: The tool shall remove backups older than the configured retention policy, shall never remove the most recent verified good backup regardless of its age, and shall provide a dry-run mode that lists exactly what would be removed and deletes nothing. Rationale: Retention code exists to delete things, which makes it the most dangerous code in this repository. Its job description and its worst failure are the same action. Acceptance criteria:

  • Given a retention policy of 7 daily generations and 14 backups present of which only the oldest is verified good, when prune runs, then the oldest is retained despite exceeding the policy, and the run record states why.
  • Given any prune invocation, when it is run in dry-run mode, then it lists every artifact it would remove with its date and size and removes nothing, and the real run removes exactly that set.

FR-OPS-01 — Every run leaves a machine-readable record

Priority: Must Requirement: Every backup, verification, restore, and prune shall append one structured record containing a run identifier, the operation, the start and end time, the bytes moved, the outcome, and — on failure — the reason; and every failing operation shall exit with a non-zero status. Rationale: The original sin of the operator’s current setup is a job that fails silently. A non-zero exit is what lets the operating system’s scheduler notice, and a structured record is what lets anything else notice. Acceptance criteria:

  • Given any of the four operations completing successfully, when it finishes, then exactly one record is appended with all required fields populated and the exit status is zero.
  • Given an operation that fails, when it finishes, then the record’s outcome is a failure with a reason naming the failing component, and the exit status is non-zero and distinct from the success status.

FR-OPS-02 — Alert on repeated failure

Priority: Should Requirement: When the configured number of consecutive backup or verification runs have failed, the tool shall emit a notification through the channel named in its configuration, and shall not emit again for the same ongoing failure until a run succeeds. Rationale: One failed night is noise. Three is an outage of your safety net. Alerting on every failure trains the operator to ignore alerts, which is worse than not alerting. Acceptance criteria:

  • Given a threshold of 3 and two consecutive failures, when the second completes, then no notification is emitted.
  • Given the same threshold and a third consecutive failure, when it completes, then exactly one notification is emitted naming the failure reason, and a fourth consecutive failure emits none.

FR-BKUP-03 — Incremental and deduplicating backups

Priority: Won’t (this release) Requirement: Block-level or file-level incremental backup with deduplication across generations. Not built. Rationale: It is the obvious next feature and it is a project of its own — a chunking scheme, an index, garbage collection of unreferenced chunks, and a restore path that reassembles across generations. Full backups of a small service are entirely adequate and are dramatically easier to verify, which is the property this tool exists to have. Revisit when the dataset makes full backups impractical, which for this operator it does not.

Starter non-functional requirements

NFR-PERF-01 — Backup and verify inside the maintenance window

Priority: Must Requirement: A full backup followed by a verification restore completes in under 10 minutes combined for a 2 GB source set, on the reference machine named in the technical specification, with the destination on local or attached storage. Measured by: three timed runs against a generated 2 GB fixture, recorded in docs/test-results.md. State the machine; a number without a machine is not a threshold.

NFR-REL-01 — The fault matrix is green

Priority: Must Requirement: Across a minimum 20-case fault matrix — destination full, destination permission removed mid-run, source database unavailable, source file deleted mid-read, process killed during write, process killed during verification, clock moved backward between runs, corrupted artifact, and a repeat of each — every case ends in either a complete verified backup or an unambiguous failure with a non-zero exit and no artifact recorded as good, and zero cases produce a partial artifact treated as complete. Measured by: an automated harness in tests/ that fails the build on any deviation. Build this in Week 3.

NFR-SEC-01 — Credentials and artifacts are not readable by others

Priority: Must Requirement: In all environments, no credential ever appears in the repository, in a process argument list, or in any log or run record; credential files and backup artifacts are created with owner-only permissions; and the tool refuses to start if a credential file is group- or world-readable. Measured by: an automated scan of logs and run records for known secret values in tests/, plus a permission assertion at startup, plus a repository history scan before your Week 8 tag. This is defensive hardening of your own service and nothing more.

NFR-A11Y-01 — Output is readable without color and by a screen reader

Priority: Should Requirement: With output piped to a file, and with output read aloud by a screen reader: no status, warning, or failure is indicated by color alone; the widely used NO_COLOR convention and a non-terminal output stream both produce plain text; every line stands alone as a sentence rather than depending on a box-drawn table; and help text is readable at 80 columns. Measured by: a snapshot test of piped output plus one recorded manual pass. A tool whose failure state is “the row turned red” is a tool that fails silently for some of its users.

NFR-OPS-01 — A stranger can restore from the runbook alone

Priority: Must Requirement: A person who is not the author performs a full restore onto a clean machine in under 45 minutes, using only docs/runbook.md and the repository, with no help from you and no undocumented step, ending in a service that passes the same structural checks used in verification. Measured by: one dated drill before Week 8, with every question the tester had recorded and fixed in the runbook. This is the course’s clean-machine test and it is also the entire point of the project.

The genuinely hard part

Consistency and partial failure, and the fact that you cannot find either by using the tool normally. A backup of a running database taken by copying its files is not a backup — it is a torn snapshot that will restore into a corrupt database, and it will look completely fine until the day you need it. Getting this right means using the database’s own dump or online-backup mechanism, or a filesystem-level snapshot, and proving it by restoring under load rather than reasoning about it. Then every one of the twenty fault cases in NFR-REL-01 is a place where a normal-looking implementation quietly does the wrong thing: a destination that fills at 60%, a process killed between writing the artifact and writing the manifest, a clock that moves backward and makes your retention arithmetic delete the wrong generation.

Which brings the honest warning about the retention code. It is thirty lines and it is the most dangerous code you will write this term, because its entire purpose is to delete backups, and a bug in it destroys exactly the thing the project exists to protect. Give it a mandatory dry run, give it a hard refusal to remove the last verified good artifact, give it its own tests before its own implementation, and — this is the part students skip — test it with the clock moved, because dates are where it will actually break.

The schedule-eater, though, is the test rig. To satisfy NFR-REL-01 you need to create a real source with real data, corrupt it in twenty specific ways, and assert the outcome, repeatably, in CI. That is real engineering and it is where a third of your construction hours go. Budget it in Week 3 and treat it as a deliverable, not overhead, because it is what makes every claim in your Week 8 talk defensible instead of anecdotal.

A stack that fits

A single command-line program in a language you already know well, plus containers for the test rig. If you know Go or Rust, you get a single self-contained binary the operator can drop on a machine, which is a genuine handoff advantage for a volunteer with no interest in managing a runtime. If Python is your fluent language, use Python and pin your dependencies — the handoff cost is real but it is smaller than the cost of learning a systems language in Week 5.

Three specific recommendations with reasons. Do not build a scheduler; use the operating system’s — cron, a systemd timer, or the platform’s task scheduler — because your tool’s job is to run correctly once and exit with a meaningful status, and a daemon is a whole second project with its own failure modes. Do not write your own database dump logic; shell out to the database’s own tool (pg_dump for PostgreSQL, SQLite’s online backup for SQLite) and check its exit status, because those tools exist precisely to solve the consistency problem you would otherwise solve badly. Keep the destination behind one small interface with a local-filesystem implementation first; if you add a remote destination later, it is one more implementation rather than a rewrite, and you avoid naming any particular provider in your requirements.

Novelty load: two unfamiliar technologies at most. Containers for the fault harness plus one new language is already two, and this brief’s difficulty is entirely in correctness — spending your unfamiliarity budget on tooling here is a bad trade. Everything above is a recommendation to adopt or deviate from in Week 3, in an ADR in docs/adr/, with the alternatives you considered. Deviate with a reason and you gain marks; deviate by default and you lose them.

What to cut first if Week 4 says you are behind

  1. Remote destinations. Local or attached storage only; document the gap honestly in the README as a known limitation. Saves 8–10 hours including all the credential handling and partial-upload resumption.
  2. Alerting (FR-OPS-02). Non-zero exit codes plus the structured run record only; the operator’s own scheduler mails the output, which is what it already does. Saves 3–4 hours.
  3. Prune (FR-RET-01) becomes a documented manual procedure with a listing command that shows what is eligible. Saves 6–7 hours, and it removes the most dangerous code in the project, which is not a bad trade under pressure.
  4. One source type instead of two. Database only, or files only, with the interface left in place for the other. Saves 6–8 hours and shrinks the verification checks correspondingly. Make this cut with your instructor; it narrows what the tool is for.

If you are ahead

Add encryption at rest with a tested key-recovery drill: artifacts encrypted before they leave the machine, and a documented procedure that restores from nothing but the archive and the key, performed once by somebody who is not you, on a clean machine. Roughly 14 hours and worth every one of them in a Week 8 demonstration — but understand what you are adding. Key management is the hard part and the dangerous part: a backup encrypted with a key nobody can find is not a backup, it is a very tidy way of deleting your data. If you add this, the key-recovery drill is not the bonus, it is the requirement, and docs/runbook.md must be able to walk a stranger through it.