Pilot Results Report
Apologetic question: "Why not despise the day of small things?"
Project 10 — Pilot Results Report
“For whoever has despised the day of small things shall rejoice, and shall see the plumb line in the hand of Zerubbabel.” — Zechariah 4:10 (ESV)
Chapter: Chapter 10 — Pilot Experiments
Due: End of Week 10 (before the full execution in Week 11)
Submit: A pilot-results-report.txt (or PDF) committed to your project’s Git repository, plus the pilot config, the completed pilot checklist, and your portfolio link.
Allowed tools: Your full research environment (Appendix A); a coding agent is allowed to write pipeline code that you verify by running small.
AI policy: Open. Use AI to accelerate the coding and the prose. You verify every number, every line of pipeline code, and every claim. Disclose substantive AI assistance in the report per Appendix C. “The model did it” is never a defense; you own every number.
The Setup
You have an environment, a design, and a budget. Your advisor — or the reviewer you’ll face in three months — does not want to see the full experiment yet. They want to see the pilot: proof that your pipeline does what you think it does, that your baselines are fair, that your splits don’t leak, and that you’ve costed the full run before spending it. This is the gate every careful researcher passes through and every burned researcher wishes they had.
The Pilot Results Report is the document that opens that gate. It is short. It is unglamorous. And it is the single best predictor of whether the work in Weeks 11–14 will stand up. A report that honestly says “NO-GO — I found a leak and a metric pointed the wrong way” is worth more than one that says “GO” without having looked.
Learning Targets
You will demonstrate that you can:
- Run your exact pipeline at pilot scale and read the small numbers correctly.
- Sanity-check a metric (range, direction, three-way agreement) and a baseline (trivial + strong, fairly run).
- Detect and rule out data leakage and a broken split on your own data.
- Quantify run-to-run variance across seeds and compare it to your effect.
- Extrapolate compute from pilot to full run and make a defensible go/no-go call.
- File the artifacts (report, config, checklist) into a reproducible portfolio with full provenance.
Normal Tier
The Pilot Results Report that meets this week’s bar.
Required features
Produce pilot-results-report.txt using code/pilot-results-report-template.txt, containing:
- What ran — method, trivial baseline, dataset/benchmark, metric + direction, split, number of seeds, compute used. One clear paragraph.
- The small numbers — a table of pilot metric (mean ± spread over N seeds) for your method, the trivial baseline, and a strong baseline.
- Sanity checks — the checklist from §10.3–10.5, each ticked with evidence: metric in range, beats trivial baseline, no split leakage, seed spread reported, loss actually moved.
- Red flags found — every surprise, what you suspected, what you confirmed.
- Compute extrapolation — pilot cost → honest full-run estimate (watch super-linear), checked against budget.
- Fix list — a numbered list of what you’ll fix before the full run.
- Go/no-go decision — explicit, with a one-line justification.
- The completed
code/neurips-style-pilot-checklist.txtcommitted alongside. - Every number carries provenance (commit hash, config, seed). AI assistance disclosed.
Normal-tier rubric (out of 100)
| Criterion | Points |
|---|---|
| Pilot actually ran your real pipeline at small scale (config committed, reproducible) | 20 |
| Metric sanity-checked: range, direction, three-way agreement documented | 15 |
| Both baselines present and the strong baseline run fairly | 15 |
| Leakage / split integrity explicitly tested and reported | 15 |
| Seed spread reported and compared to the effect | 10 |
| Compute extrapolation and budget check | 10 |
| Honest go/no-go decision with justification | 10 |
| Provenance + AI disclosure + portfolio filing | 5 |
| Total | 100 |
Medium Tier (+up to 25% extra credit)
Stronger rigor. Add any of:
- Two leakage classes audited, not one: exact-duplicate detection across splits and (for an LLM project) a benchmark-contamination check against a contamination-resistant benchmark (LiveCodeBench / LiveBench / MMLU-Pro / FrontierMath, as of 2026), with the risk stated honestly.
- Overfit-a-batch evidence included as a figure (training loss → 0 on 16 examples), proving the wiring before the data question.
- A second metric piloted to confirm your conclusion isn’t an artifact of one metric’s quirks.
- A cost-aware scale-down plan: if the full run exceeds budget, a concrete revised plan (fewer seeds, smaller grid, a cheaper rung on the Prompt→RAG→Fine-tune ladder) with its own re-estimate.
Hard Tier (+up to 25% additional extra credit)
Publication-grade judgment — the part a tool cannot do for you.
Write a one-page Pilot Decision Memo addressed to your advisor that argues a non-obvious call. Not “the checks passed, GO,” but a reasoned judgment under uncertainty: e.g., “the pilot effect (1.2 pts) is barely above the seed spread (0.9 pts); I recommend a GO conditional on doubling seeds and adding the ablation, because the cost of being wrong here is a wasted full run, but the cost of abandoning is losing the only novel angle in my lit-review gap.” The memo must weigh the cost of a false GO against the cost of a false NO-GO for your specific project and budget, name the threats to validity (Chapter 6) the pilot could and could not address, and commit to a decision. An AI can summarize your numbers; it cannot own this decision. You can.
Submission
Commit to your project repository:
pilot-results-report.txt(or PDF)configs/pilot.yaml(the pilot configuration)- the completed pilot checklist
- (Medium/Hard) the figures and the decision memo
Submit the repo link + the commit hash of your pilot. Confirm the report is filed in your research portfolio alongside your experimental notebooks and draft manuscript.
Hints
- Start with Rep 2 (overfit 16 examples). If that fails, nothing downstream matters yet — fix the wiring first.
- The Pilot Sanity Checker widget (§10.8) is the fastest way to train your red-flag reflex before you trust your own pilot.
- A 0.99 on the pilot is almost always a bug. Treat “too good” as a trigger for more scrutiny.
- Don’t tune until the number looks good. That’s p-hacking (Chapter 12), and the pilot is not the place to chase a result.
- If you genuinely can’t run the full pipeline yet, write the NO-GO honestly — that is a passing, valuable report.
What Mastery Looks Like
A reader of your report can rerun the pilot from your repo, agrees with your go/no-go call, and finds nothing you hid. Your “too good” numbers were investigated, not celebrated. Your strong baseline was run as carefully as your own method. Your effect is reported next to its seed spread, not alone. And your fix list reads like the to-do list of someone who intends to publish something true.
Coach’s Note — The grade rewards the quality of your suspicion, not the size of your pilot number. A NO-GO that caught a real leak outscores a GO that looked the other way. Be the researcher who is glad the pilot broke.
When You’re Done
- Pilot ran on the real pipeline at ~10% scale, config committed.
- Metric checked three ways; range and direction confirmed.
- Trivial + strong baselines run; strong one run fairly.
- Leakage / split integrity explicitly tested.
- Seed spread reported next to the effect.
- Compute extrapolated and checked against budget.
- Go/no-go decision made and justified.
- Checklist completed; provenance on every number; AI disclosed; filed in portfolio.
A theological footnote. Zechariah’s word came to people whose new foundation looked like nothing next to the temple they remembered, and the word was: do not despise it. The plumb line in Zerubbabel’s hand is the small, exact tool that makes the great building true. Your pilot is that plumb line — the small, exact, unglamorous check that keeps the whole work honest. To run it carefully, to let it correct you, to bear true witness to what the small numbers say even when they embarrass you — this is faithfulness in little things, and it is the same faithfulness the great work will require. Do not despise the day of small things.