Project 6

Question, Hypothesis, and Variables

Apologetic question: "What does it mean to make a claim you could be wrong about?"

Project 6 — Question, Hypothesis, and Variables

“The one who states his case first seems right, until the other comes and examines him.” — Proverbs 18:17 (ESV)

Chapter: 6 — Research Questions, Hypotheses, and Variables Due: End of Week 6 (see the course schedule). Submit: A question-hypothesis-variables.txt document committed to your portfolio repo, plus the filled code/variables-worksheet.csv and your threats table. Tag the commit p06-submit. Allowed tools: Anything — Zotero, the comparison matrix, AI assistants. This is open-AI. AI policy: You may use AI to enumerate rival hypotheses, candidate confounds, and metric failure modes, and to stress-test wording. You may not let it choose your metric or author your hypothesis. Disclose substantive AI use in an AI-use note per Appendix C. The human owns every claim.


The Setup

You have a domain, a gap, and an AI method you think advances that domain. What you do not yet have is a claim a reviewer could attack — and that is the only thing the rest of the practicum can be built on.

This week you write the methodological core of your eventual paper: a research question precise enough to answer, a hypothesis sharp enough to lose, the variables you will vary and measure and control, and an honest accounting of the four ways you might be fooling yourself. This is small in page count and enormous in leverage. Every later deliverable inherits its quality from this one. The Research Proposal (P8, 15%) is largely an expansion of this document; the experimental-design document (P7) executes it; the results, discussion, and paper all answer the claim you fix here.

Treat it as the load-bearing wall it is. A crack here is invisible this week and structural by Week 14.


Learning Targets

You will demonstrate that you can:

  • Climb the topic → question → hypothesis → operationalized-claim ladder for your own work.
  • Write a falsifiable hypothesis with an explicit null and a predicted direction of effect.
  • Specify independent, dependent (primary + secondary), and controlled variables, each operationalized.
  • Distinguish a construct from its proxy and name the gap between them.
  • Identify and mitigate the four Cook & Campbell threats to validity for your specific design.
  • Make — and state — the deliberate internal-vs-external validity tradeoff your design implies.

Normal Tier

Meets the deliverable bar for Week 6: a complete, falsifiable, operationalized claim for your domain with a four-category threats table.

Required features

  1. Research question — one sharpened question, traceable to your Week 4 gap, containing a comparison.
  2. Hypothesis + null — H₁ stated with predicted direction; H₀ stated cleanly. No hedge words (“can,” “may,” “tends to,” “sometimes”). Verified clean by the Hypothesis & Validity Builder.
  3. Variables — exactly one IV; a primary DV and at least one secondary DV (a tradeoff); at least three controlled variables (base model + version/date, data + split, compute). Each operationalized. Delivered in code/variables-worksheet.csv.
  4. Metric operationalization — construct / proxy / gap / direction-of-error for your primary DV.
  5. Threats-to-validity table — at least one concrete, design-specific threat per Cook & Campbell category, each with a concrete mitigation. Use code/threats-to-validity.txt.
  6. Seeds + test discipline — a stated number of seeds (≥ 5) and the test split you will touch only once.

Normal-tier rubric (out of 100)

CriterionPoints
Research question is precise, comparative, and traceable to the Week 4 gap15
Hypothesis is falsifiable; null stated; no hedge words; direction predicted25
Variables: one IV, primary + secondary DV, ≥3 controls, all operationalized20
Metric: construct/proxy/gap/direction-of-error named honestly15
Threats table: ≥1 concrete threat + mitigation per all four categories20
Seeds (≥5) and single-touch test split stated5
Total100

Medium Tier (+up to 25% extra credit)

Stronger rigor. Pick up to all of:

  • Factorial honesty (+8%). If your real question has two factors (e.g., retrieval × model size), specify a proper 2×2 factorial design with the interaction term named — instead of pretending it’s one IV.
  • Rival-hypothesis register (+8%). Maintain a short list of competing explanations for your predicted result and the discriminating measurement that would separate them. (Use AI to enumerate; you decide which are live.)
  • Pre-registration draft (+9%). Draft your design as an OSF or AsPredicted-style pre-registration (RQ, hypotheses, variables, analysis plan) — time-stamped before any data. This is the strongest single defense against HARKing and p-hacking, and it sets you up for Week 8.

Hard Tier (+up to 25% additional extra credit)

Publication-ready judgment a tool cannot supply. Write a 1–2 page validity memo that does what no AI can do for you:

  • Argue, in prose, the single hardest threat to your construct validity — the case that your metric does not measure what you claim — and defend (or honestly concede) it. Engage the live applied-AI debate directly: if your DV rests on a benchmark, address contamination and the “does this score measure the construct or memorization?” problem (the MMLU-reasoning-vs-memorization question from §6.5).
  • Make the internal-vs-external tradeoff decision explicit: state where your design sits, what you are buying, what you are giving up, and why that is the right call for your research question — not in general.
  • Name the result you most fear, and commit, in writing, to reporting it if it happens.

The memo is graded on the quality of the judgment, not the polish. A confident hand-wave scores low; an honest “here is exactly where I am vulnerable and here is why I proceed anyway” scores high.


Submission

Commit to your portfolio repo and tag p06-submit:

  • question-hypothesis-variables.txt — the four artifacts assembled.
  • code/variables-worksheet.csv — your filled variables table.
  • The threats-to-validity table (in the main doc or as threats-to-validity.txt).
  • AI-use.txt — any substantive AI assistance, per Appendix C.
  • (Medium/Hard) the pre-registration draft and/or validity memo.

Drop the variable and threat definitions into your acmart/IEEEtran methodology section now (Appendix B) — the practicum is a submission, not a class project.


Hints

  • Write the null first if you’re stuck. “No difference in X between A and B” is mechanical; once you have it, H₁ is just “and I predict A > B.”
  • One IV. If you have two, you have a factorial design (Medium tier) or a confound (a mistake). Choose deliberately.
  • The secondary DV is not optional. A method with no measured cost is a sales pitch. Find the tradeoff.
  • Run the widget early and often. It catches hedge words and surfaces threats faster than you can.
  • Borrow structure from a real paper. Rep 10 in the exercises had you dissect one — model your threats table on a published one in your area.

What Mastery Looks Like

A masterful submission reads like the methodology section of a paper that a reviewer would trust. The hypothesis sticks its neck out. The null is clean. Every variable is operationalized so precisely that a stranger could rebuild the experiment. The threats table names the real weaknesses — including the one in the category you’re weakest in — and the mitigations are specific actions, not reassurances. And the Hard-tier memo demonstrates the rarest thing: a researcher who examined their own case before anyone else could.

Coach’s Note — The temptation this week is to make the claim safe so it can’t be attacked. Resist it completely. A claim that can’t be attacked also can’t be confirmed — it teaches nothing. The bravery of a sharp, losable hypothesis is the whole point. Write the sentence that could be wrong.


When You’re Done

  • Research question is comparative, precise, traceable to the Week 4 gap.
  • Hypothesis is falsifiable, hedge-free, with predicted direction; null is stated.
  • One IV; primary + secondary DV; ≥3 controls; all operationalized.
  • Metric: construct / proxy / gap / direction-of-error named.
  • Threats table covers all four Cook & Campbell categories with concrete mitigations.
  • Seeds (≥5) and single-touch test split committed.
  • AI use disclosed; committed and tagged p06-submit.

A theological footnote. This week asks you to write a claim you could be wrong about — and to be the first to examine it. “The one who states his case first seems right, until the other comes and examines him” (Prov 18:17, ESV). The discipline of falsifiability, of the honest null, of the self-named threat, is the practice of refusing to merely seem right. It is a small humility built into the structure of the method: a confession that truth is not ours to author, only to serve, and that conforming our claim to reality matters more than protecting our reputation. The researcher who cross-examines their own hypothesis before a reviewer does has already chosen truth over the appearance of it. Do the work heartily, as for the Lord — which here means doing it honestly enough to be proven wrong.