Project 13

Discussion Section

Apologetic question: "What does it mean to test everything and hold fast what is good?"

Project 13 — Discussion Section

“but test everything; hold fast what is good.” — 1 Thessalonians 5:21 (ESV)

Chapter: Chapter 13 — Discussion: Making Meaning Due: end of Week 13 (before the Week 14 draft-assembly lab) Submit: a discussion.pdf (compiled from your acmart/IEEEtran Overleaf project) plus your exported claim-evidence map and your updated comparison matrix, committed to your portfolio repo and linked. Allowed tools: your full research stack — Overleaf, Python/SciPy/deep-significance, Zotero, dblp/Semantic Scholar for citation verification. AI policy: open-AI for phrasing and structure; never for claims. Every claim-bearing sentence must pass a claim-evidence map against a real result before it stays; every cited comparison must be verified at its source. Disclose any substantive AI assistance in your Acknowledgements per your target venue (Appendix C). The model did not run your experiment — you did, and you own every sentence.


The Setup

You are three weeks from a draft paper. Your raw results are in (Chapter 11); your analysis is honest (Chapter 12). Now a Program Committee member — the imaginary one who will decide your paper’s fate — opens your Discussion section first, because that is where they learn whether to trust you. They are not reading for your enthusiasm. They are reading, sentence by sentence, for the gap between what you claim and what you measured.

Your job this week is to write the Discussion that committee member cannot reject for overclaiming. Not a hype piece. Not a hedge so timid it says nothing. A precise, defensible interpretation: here is what we found, here is how it sits against prior work, here is exactly how far it generalizes, and here is where it stops. The kind of section a skeptic finishes and thinks, I believe this — and I know what I’d build next.


Learning Targets

You will demonstrate that you can:

  • Interpret results in plain language tied to your research question, without restating the Results table.
  • Compare your findings against prior work explicitly, using your comparison matrix and naming any asymmetries (e.g., single-run vs. multi-seed).
  • Distinguish statistical from practical significance, leading with effect size and CI (ASA 2019), not bare p-values.
  • Handle unexpected and negative results honestly, separating preregistered question from post-hoc hypothesis (no HARKing).
  • Bound the conclusion with one specific limitation per Cook & Campbell validity category, including a construct-validity reckoning with benchmark contamination where relevant.
  • Map every claim to its specific evidence, leaving zero unsupported claims.

Normal Tier

Required features

  1. A Discussion section (~600–1000 words) in your acmart (sigconf) or IEEEtran Overleaf project, compiling to PDF.
  2. Interpretation of your headline result, tied explicitly to your research question — not a restatement of the table.
  3. Comparison against at least three prior-work rows from your updated comparison matrix (code/comparison-matrix.csv), naming at least one asymmetry in the comparison.
  4. Significance handled honestly: your headline claim stated with effect size and confidence interval, not the p-value alone.
  5. Honest treatment of at least one unexpected, negative, or null result from your runs.
  6. A limitations paragraph with one concrete limitation per Cook & Campbell category (statistical conclusion, internal, construct, external).
  7. A claim-evidence map (code/claim-evidence-map.txt or the widget export) with zero Unsupported claims and every Over-reach narrowed.
  8. Every cited comparison verified at dblp or Semantic Scholar; AI assistance disclosed per Appendix C if used.

Normal-tier rubric (out of 100)

CriterionPoints
Interpretation tied to the research question (not table-restatement)18
Explicit comparison to prior work + named asymmetry16
Effect size + CI lead; statistical-vs-practical distinction honest14
Honest handling of an unexpected/negative/null result12
Limitations: one specific threat per validity category14
Claim-evidence map complete; zero Unsupported, Over-reaches narrowed16
Citations verified at source; AI disclosure; clean compile10
Total100

Medium Tier (+up to 25% extra credit)

Stronger rigor and novelty:

  • Run a proper multi-seed significance comparison with deep-significance (Almost Stochastic Order, τ = 0.2 threshold; v1.2.5, arXiv:2204.06815) on your seed scores, and interpret the ASO result in the Discussion rather than a t-test.
  • Address benchmark contamination as an explicit construct-validity threat: state your model version and date, reason about whether the test set could plausibly be in training data, and bound your “better than” claim accordingly.
  • Turn at least two stated limitations into specific, scoped future-work directions — not “more experiments,” but “test transfer to language X using benchmark Y.”

Hard Tier (+up to 25% additional extra credit)

This tier demands judgment a tool cannot supply. Write a one-page interpretation memo (separate from the section, in your portfolio) that argues:

  • What your central finding licenses a reader to conclude — and what it does not. Draw the line explicitly. The most defensible memos are more restrained than the author’s instinct.
  • Why a careful skeptic should believe it anyway. Marshal your evidence — multi-seed variance, the CI, the controlled threats — into a case that survives a hostile read.
  • The one experiment that would most change your conclusion, and what you predict it would show.

An LLM can summarize your tables. It cannot decide where the truth of your claim ends, because it did not run your experiment and bears no accountability for the conclusion. This memo is you doing the irreducibly human part of science: weighing the evidence and standing behind a judgment. Grade weight here is on the reasoning, not the prose.


Submission

  • discussion.pdf compiled from your Overleaf project (note the template: acmart/sigconf or IEEEtran).
  • claim-evidence-map.txt (or widget export) showing zero Unsupported claims.
  • Updated comparison-matrix.csv with your work as a row.
  • Hard tier: interpretation-memo.pdf (one page).
  • A one-line AI-disclosure note in your README per Appendix C.
  • All committed to your portfolio repo with this week’s tag.

Hints

  • Write the limitations paragraph first. It is the easiest to write honestly and it forces you to know your conclusion’s edges before you write the bolder interpretation sentences.
  • When a claim has no table to point at, you have two honest moves and no third: measure it (and move the number to Results), or delete the sentence. “Trust me” is not a move.
  • Read your three most-cited prior-work papers’ Discussion sections for structure, not content — notice how they bound their claims, then do the same with yours.
  • For the significance forms, code/three_ways.py prints all three sentences from your means and SD; copy the most honest into the section.

What Mastery Looks Like

A reviewer reads your Discussion and never once reaches for the comment “overclaims relative to the evidence.” Every sentence has a result behind it. The comparison is fair, including where it is uncomfortable. The negative result is in there, plainly. The limitations read like a map a successor could follow, not a confession. And the whole thing feels confident — not because it claims a lot, but because it claims exactly what it can defend, and not one inch more.

Coach’s Note — The temptation all week is to make your finding sound bigger than it is, because you worked hard and you want it to matter. Resist. The paper that claims a small thing it can prove beats the paper that claims a big thing it can’t, every single time, with every reviewer who has ever lived. Smaller and true wins.

When You’re Done

  • Discussion section compiles cleanly in your Overleaf template.
  • Headline result stated with effect size + CI, not bare p.
  • At least three prior-work comparisons, one named asymmetry.
  • At least one unexpected/negative result discussed honestly.
  • One specific limitation per Cook & Campbell category.
  • Claim-evidence map: zero Unsupported, all Over-reaches narrowed.
  • Every cited comparison verified at dblp/Semantic Scholar.
  • AI assistance disclosed per Appendix C.

A theological footnote. “Test everything; hold fast what is good” (1 Thess 5:21, ESV) puts the order beyond negotiation: the test comes before the holding. To write a Discussion is to bear witness — and Scripture is unambiguous that a false witness wounds a neighbor (Exod 20:16). Your neighbor here is concrete: the researcher who will read your paper and build on your claim. To inflate a finding is to send them down a road you know is weaker than you said. To bury a real finding because it is awkward is its own failure of nerve. Honest interpretation — confident about what the data say, restrained about what you wish they said — is not merely good method. It is loving your neighbor with your evidence. “The glory of kings is to search things out” (Prov 25:2, ESV); the glory is only in the searching if the finding is true.