Understanding AI's Limits
What does it mean to test the spirits — to refuse to believe every confident claim?
Chapter 8 — Understanding AI’s Limits
“Trust, but verify.” — Russian proverb
“Beloved, do not believe every spirit, but test the spirits to see whether they are from God, for many false prophets have gone out into the world.” — 1 John 4:1 (ESV)
Why This Matters
You are at the midpoint. For seven weeks you have been learning to use AI — to pick a model, to prompt like a professional, to choose the right tool, to draft an email and summarize a report. You have gotten good at it. Fast, even. This week I take the tool out of your hands, turn it over, and show you the edges.
Because here is the thing nobody selling you a subscription wants to say plainly: the most powerful professional tool of your lifetime is also confidently, fluently, and regularly wrong. Not occasionally. Not only on hard questions. It will invent a study that does not exist, attach a real professor’s name to a book she never wrote, quote a CEO saying something he never said — and it will do all of it in clean, confident, well-formatted prose that looks exactly like the truth. The wrong answers do not arrive wearing a warning label. They arrive in the same calm voice as the right ones. That is the whole problem, and this chapter is about surviving it.
This is Week 8, which means two things at once. It is the midterm — everything from Weeks 1 through 8 comes due. And it is the hinge of the book. Phase 1 was understand and choose. Phase 2, starting next week, is apply and integrate — image, video, voice, agents, real workflows. You do not get to build on top of a tool you cannot verify. So before we hand you more power, we install the brake. This week is the brake.
Everything here serves one sentence, and I want it burned in by Friday: AI is an assistant, not an authority. An assistant drafts, suggests, accelerates, and hands you something to check. An authority you obey. The difference between a professional who thrives with AI and one who gets humiliated by it is not who prompts better. It is who verifies — who treats a fluent answer as a claim to be tested, not a fact to be repeated. Say it with me, in the framing that runs through every chapter of this book:
You choose the tool. You own the verdict. AI drafts, generates, and accelerates; you decide, verify, and are accountable. When the briefing has your name on the slide, “the AI said so” is not a defense. It is a confession.
And that lands us on this week’s question — the one we will earn, not preach. In Week 7 you met the noble Bereans, who examined the claims to see whether they were so. This week the Apostle John hands you the posture underneath that habit: do not believe every spirit, but test the spirits. Not believe none — that is cynicism, and cynics build nothing. Not believe all — that is credulity, and credulous people get fooled. A third way: test. So the question for Week 8 is this: what does it mean to test the spirits — to refuse to believe every confident claim just because it is fluent? Hold it. We will earn it by the end.
Coach’s Note — If you are new to the house style: learning is a sport, the keyboard is the gym, and reading about verification makes you exactly as skilled at it as reading about push-ups makes you strong. This chapter ends with Reps, an Interactive Lab, and the midterm project. The lesson does not land until it is in your hands.
8.1 — Why a Tool This Good Is Also This Wrong
To catch the failures you first have to understand why they happen — because a limit you understand is a limit you can watch for, and a limit you think is a bug is one you will keep being surprised by.
Cast your mind back to Week 2. We said a large language model is, at its core, a machine that predicts the next word. You type; it computes, one token at a time, the most statistically likely continuation of the text so far, given everything it absorbed in training. That is astonishing and it is all it is. It is not looking anything up. It is not consulting a database of true facts. It has no internal ledger marked “things that are actually true.” It is producing the words that tend to follow the words you gave it.
Sit with the consequence. When the true answer and a plausible-sounding false answer are equally likely continuations of your prompt, the model has no built-in reason to prefer the true one. It reaches for fluent, not correct, because fluent is what it was trained to produce. Most of the time fluent and correct overlap — the internet it learned from is mostly, roughly right about common things — and so the tool feels reliable. But at the edges, where the true fact is rare or specific or recent, the model will happily generate the shape of a right answer with a wrong fact inside it. A citation-shaped string. A quote-shaped sentence. A statistic-shaped number.
This is not a defect some future version patches out. As of mid-2026, every frontier model — Claude Opus 5, GPT-5.x, Gemini 3 Pro, all of them — still does this. And reaching those same models through the vetted institutional platform Concordia licenses — BoodleBox — changes none of it: a tool that is safer for your data is not a tool that is more truthful, so you verify its output exactly like any other. No platform is exempt from the check. Newer models do it less, and that is genuine progress, but “less” is a trap of its own: a tool that is wrong 2% of the time lulls you into skipping the check on the 2%. The limitation is a property of the technology, not a phase it is growing out of. Your job is not to wait for a model that never lies. Your job is to build the verification habit that makes the model’s lying survivable.
8.2 — Hallucination: The Confident Wrong Answer
The industry’s word for it is hallucination: a confident, fluent, plausible output that is simply false. (It is in your glossary — Appendix D.) It is a soft word for a sharp problem. The model is not “confused” and it is not “guessing” in the way a nervous student guesses. It is generating with full fluency an answer that has no basis in reality, and it cannot tell you it is doing so, because it does not know.
Hallucinations come in a few recognizable species, and naming them is half the battle:
| Species | What it looks like | Where it bites |
|---|---|---|
| Invented fact | A confident claim, oddly specific, unsourceable | History, science, “how many,” “when” |
| Fabricated citation | A study/report/case/URL that does not exist | Research, legal, medical, anything “with sources” |
| Fake quote | Real person, words they never said | Speeches, “as X famously said,” marketing |
| Made-up detail | Real thing, wrong specifics (a feature, a clause, a step) | Product specs, procedures, contracts |
| Confident non-answer | A fluent essay that never actually answers | Vague or unanswerable questions |
We ship a live specimen in this chapter’s code/ folder: code/ai-answer-with-errors.txt. It is a real-shaped AI briefing on “AI adoption in small business,” the kind a coordinator might paste into a slide before a Thursday leadership meeting. It reads beautifully. It also contains a fabricated McKinsey study, a quote Satya Nadella never gave, an invented federal law, a book by an author who does not exist, an out-of-date “most capable model,” and a flatly false claim about how these tools connect to the internet. Open it. Read it cold. Notice that nothing about the prose warns you. That is the felt lesson of the whole week, and it is why the midterm exists.
Coach’s Note — The dangerous hallucinations are never the absurd ones. “The moon is made of cheese” fools no one. The one that costs you your credibility is the plausible one: a 2024 report title that sounds exactly like a report McKinsey would publish, a statistic in the range you already half-believed. Hallucinations hide inside your expectations. The more a claim flatters what you already think, the harder you should check it.
8.3 — Fluency Is Not Accuracy: Confidence vs Correctness
Here is the single most important sentence in this chapter, and it is worth reading twice: in an AI answer, confidence and correctness are unrelated.
A human expert usually signals uncertainty. She hedges, slows down, says “I think,” “I’d have to check,” “off the top of my head.” We have spent our whole lives reading those signals — a confident tone from a competent person is decent evidence the thing is true. An LLM breaks that lifelong instinct. It writes the true answer and the false answer in the identical register: same crisp confidence, same clean formatting, same authoritative cadence. The fake McKinsey statistic in our sample file is delivered with exactly as much poise as any real fact around it. The tone tells you nothing. You have to unlearn a lifetime of trusting it.
Researchers call the ideal calibration — a well-calibrated source is unsure when it should be unsure. Current models are poorly calibrated in the way that hurts you most: they are often most fluent precisely where they are making things up, because a fabricated citation has a very clean, very predictable shape to generate. Fluency, in other words, is sometimes a warning sign, not a comfort.
So retrain the instinct. When an AI answer is smooth, complete, confident, and conveniently exactly what you wanted — that is not the moment to relax. That is the moment to check. The polish is not evidence. Sometimes it is the disguise.
Coach’s Note — Try this once and you will never forget it: ask a model a question about something that does not exist. “Summarize the plot of the 1997 film The Glass Cartographer.” Many models will cheerfully invent a cast, a director, and a three-act plot for a movie that was never made — with total confidence. The confidence was never connected to the truth. It never was.
8.4 — The Frozen Clock: Training Cutoffs and Outdated Knowledge
Every model has a training cutoff — a date after which it learned nothing. Its knowledge is a photograph, not a live feed. Ask it about anything that happened, changed, or was released after that date, and one of two things happens: it tells you it does not know (the honest failure), or — far more often and far more dangerously — it confidently answers as if the world stopped on its cutoff day.
Our sample file has a textbook case: “the most capable model on the market today is OpenAI’s GPT-4, launched in March 2023.” Every word of the launch fact is true. The framing — today, most capable — is badly false, because the model’s “today” is frozen years in the past. This is the sneakiest error class, because it is built from real facts wearing the wrong tense.
Two complications make this worse in 2026, and you need both in your head:
- Some tools now browse the live web; most base chats still do not — by default. A research mode or an “answer engine” like Perplexity can fetch current pages — and in BoodleBox you reach exactly that by opening the bot picker and @-mentioning a browsing model like Perplexity in the chat (as of 2026), instead of trusting a base model’s frozen answer. But the default chat box in a consumer app often answers from frozen training weights unless it chooses to search, and it does not always tell you which it did. Our sample file’s line — “Because these models are connected to the live internet by default, their answers always reflect the latest information” — is simply false as a general claim, and it is exactly the kind of thing a professional gets burned by.
- Prices, versions, laws, and “latest” claims rot fastest. Anything with a version number or a dollar sign is a landmine. This entire book pins model IDs and dates every price as a “mid-2026 snapshot” for exactly this reason. When an AI gives you a price, a version, or a “current best,” treat it as a lead to verify on the vendor’s own page — never as a fact.
The habit: whenever the answer depends on recency — a price, a release, a law, a “who is winning” — assume the model’s clock is stopped and go check a live, dated, trusted source yourself.
8.5 — Bias: The Model Learned From Us
An LLM learns from an enormous slice of human writing, and human writing carries human bias. So the model absorbs it — statistical patterns about who does which jobs, which names go with which places, which viewpoints show up as “default” and which as “other.” Labs work hard to reduce the ugliest of this, and they have made real progress. But bias is not a switch you flip off. It is baked into the training data, and it surfaces in quiet, professional-looking ways:
- Representation bias — ask for “a picture of a nurse” or “a CEO” and watch which demographics show up by default. (We will feel this hard in Week 9’s image work.)
- Framing bias — the model presents one side of a genuinely contested question as settled fact, because that framing was more common in its training data.
- Cultural default bias — it assumes U.S. norms, English idioms, Western holidays, and one set of professional conventions unless you tell it otherwise.
- Sycophancy — a subtler bias: models tend to agree with you. Push back and many will fold and reverse a correct answer, because agreeable text is what they were rewarded for. If you can talk a model out of the truth, you have learned something about the tool, not the truth.
For a working professional, the danger is not usually a cartoonish slur. It is a summary, a shortlist, or a “neutral” description that quietly skews — a hiring blurb, a market analysis, a “typical customer” persona that flattens real people into a stereotype. The fix is the same as everywhere else: you are the editor. Read for whose perspective is missing. Ask “who does this leave out?” and “would I present this framing as the only one?” The tool has a worldview it did not choose. You have a responsibility it cannot carry.
8.6 — Fabricated Citations and Fake Quotes: The Failure That Ends Careers
If I could make you paranoid about exactly one thing, it is this: AI invents sources, and sources are what professionals stake their credibility on.
The most-cited cautionary tale is real and worth knowing. In 2023, two New York attorneys filed a legal brief that cited a series of court decisions to support their case. The cases did not exist. ChatGPT had fabricated them — names, quotes, citation numbers, all of it — and the lawyers, not knowing the tool could do that, filed them. When opposing counsel could not find the cases, the truth came out; the court sanctioned the attorneys (a fine of around $5,000, widely reported) and the story went around the world. The tool did not fail them. Their verification failed them. They treated a fabricated citation as a real one because it had the shape of a real one.
This failure has a signature you can learn:
- A perfect-looking citation that returns nothing. Author, year, journal, a plausible title — and when you actually search for it, silence. If you cannot find the source, assume there is no source.
- A real source that says something else. Even scarier: the paper exists, the model just misrepresents what it found. You have to open it, not just confirm it exists.
- “Deep research” is not a free pass. The research and answer engines that browse and cite are a genuine step up — but their citations can still be wrong, mismatched, or invented. A footnote is a promise, not a proof. Click it.
- Fake quotes. “As [famous person] said…” is a red flag, not a credential. Our sample file’s Nadella-at-Davos quote is fabricated. Real quotes have a findable primary source — a transcript, a recording, a verifiable article. If you cannot find one, do not repeat it under your name.
The rule is short: a citation you did not open is not a citation. It is a decoration.
Coach’s Note — This is the failure I most want to keep off your record. Everything else on this list embarrasses you. Fabricated sources can get you sanctioned, fired, retracted, or sued, because you did not just repeat a wrong fact — you claimed an authority for it that never existed. Open every source. Every single one. There is no faster way to end a professional’s credibility than one made-up citation caught by someone else.
8.7 — Your Verification Toolkit: Five Moves
Enough diagnosis. Here is the treatment — the five moves that turn “the AI said so” into “I checked, and here is what is true.” We ship them as a printable playbook, code/verification-checklist.txt; this is the narrated version.
- Cross-check. Confirm the claim in at least one independent, trusted source you did not get from the AI. Independent is the key word — a second AI that learned from the same internet is not independent.
- Ask for the source — then follow it. Make the model cite its claim, then open the source. Half the time the “source” evaporates when you click it. Asking for citations is not verification. Reading them is.
- Check against a trusted anchor. For facts that matter, go to the primary source: the official page, the vendor’s own pricing, the actual study, the government site, a real subject-matter expert. Anchor to something you trust more than the AI.
- Triangulate across models. Ask the same question of a second, different model (this is exactly the muscle you built in Week 2’s three-model comparison). In BoodleBox you no longer need a second tab: type
@to open the bot picker and pull another model into the same chat — a multi-bot chat, as of 2026 — then read the two answers side by side. Where they disagree is a flare marking the spot to dig. But hear the caveat: agreement is not proof. Two models trained on the same wrong data will confidently agree on the same wrong answer, and a shared platform does not make them independent. (Off campus, a second free tool in a second tab does the same job.) - The stake-your-name test. The last gate, and the one that makes you a professional: Would I put my name under this — in front of my boss, my client, a regulator — and answer for it if it turns out wrong? If the honest answer is no, it is not verified. It is a rumor with good grammar. Do not ship it.
You will not run all five on every sentence — that would be paralysis, and paralysis is its own failure. You run them in proportion to the stakes. A brainstorm needs a glance. A number on a board slide needs move 3. A citation in a report that goes to a client needs moves 2, 3, and 5, every time. Which brings us to the gate that decides how hard to check.
8.8 — When Not to Use AI
Appropriate trust starts with a prior question: should the machine be anywhere near this decision at all? Not every task belongs to AI, and the professional’s real skill is knowing which ones do not. Run everything through a two-question gate:
- Q1 — How high are the stakes if this is wrong? Money, health, safety, legal exposure, someone’s reputation, a decision that cannot be undone.
- Q2 — How easily can I verify it against a trusted source?
| Easy to verify | Hard to verify | |
|---|---|---|
| Low stakes | Green. Ideal AI work — draft freely (brainstorm a subject line). | Yellow. Fine for a first draft; label it “unverified.” |
| High stakes | Amber. AI may draft; you verify every line before it ships (a client email with a figure in it). | Red. Do not rely on AI. Use a human expert or the primary source (a medical dose, a legal filing, a safety procedure, a diagnosis). |
The red box is not a failure of AI; it is a boundary of it. Some judgments are yours because the stakes are real, the answer is hard to verify, and a human has to be accountable. A nurse does not confirm a dosage with a chatbot. A lawyer does not file what she has not read. You do not send the board a number you cannot source. And a rule that spans every box: never paste confidential information, client data, PII, or health records into a consumer AI tool — that is not a verification question, it is a trust-and-law question, and it lives in Appendix C.
Coach’s Note — The mark of maturity with this technology is not using it for everything. It is knowing the handful of decisions you will never hand to it — and being able to say why out loud. The person who uses AI for nothing is a Luddite. The person who uses it for everything is a liability. The professional knows the line and can defend it.
8.9 — Appropriate Trust: Calibration, Not Cynicism
I have spent a whole chapter showing you how the tool fails, so let me close the loop before you overcorrect. The goal of this week is not to make you distrust AI. A person who trusts nothing gets none of the enormous value on the table — and there is enormous value, or this book would not exist. The goal is calibration: trust that matches reality, dialed to the stakes.
Think of a trust dial, not a trust switch. Low-stakes, easily-verified, creative, drafting work — turn it up; let the tool run. High-stakes, hard-to-verify, consequential work — turn it down; verify everything, or keep the tool out entirely. The skill is not a fixed setting. It is reading each task and setting the dial. That reading is critical thinking, and critical thinking is the one thing on this whole list AI cannot do for you, because it is the act of deciding how much to trust the AI.
This is what “human in the loop” actually means, stripped of the jargon. It does not mean a human is nearby. It means a human is accountable — that at the point where judgment lives, a named person decided, verified, and will answer for the outcome. The AI compressed the work of getting to the decision. It did not take the decision, and it did not take the accountability, because a machine cannot be held to account. Only you can. That is not a limitation to lament. It is the exact place your professional value now lives.
8.x — Interactive Lab: Hallucination Hunt
Below this chapter on the website you will find an interactive panel called the Hallucination Hunt. Go use it now — it is not decoration, it is the rep that wires this whole chapter into your hands.
The Hunt drops you a fluent, confident AI answer — the kind you would happily paste into a slide — and gives you one job: click every sentence that contains a fabricated fact, an invented citation, a fake quote, an outdated claim, or a plausible-wrong number. When you commit your selections, the Hunt scores you against ground truth and, for each one you flagged (and each one you missed), reveals why it is wrong and which species of error it is. You will miss some. Almost everyone does on the first pass — and that miss, felt in your own hands, teaches more than any paragraph I could write.
Run it twice. The first pass, go cold — no notes, no checklist, just your eye. Record your score; it will humble you, and it is meant to. Then open code/verification-checklist.txt, study the six error tells, and run the Hunt again — this time hunting deliberately by type. Watch your score climb. That climb is the skill: verification is not a talent you have or lack, it is a discipline you can install. The Hunt is where you install it before it counts on the midterm.
In BoodleBox — The Hunt scores you on a frozen answer; your real answers come out of a live tool, and the same discipline points straight at it. When an answer actually matters, don’t check it alone: in BoodleBox, type
@to open the bot picker and pull a second model into the same chat (a multi-bot chat, as of 2026), ask both the same question, and read where they disagree — that gap is your dig site. Sign in with your Concordia account at box.boodle.ai; any second free tool works the same way off campus. And carry 8.1 with you: BoodleBox is the safe place for your school and work information, not a truthful one — its models hallucinate too, so verify their output like any other.
8.10 — Test the Spirits
Now the week’s question, given its full weight. What does it mean to test the spirits — to refuse to believe every confident claim just because it is fluent?
The Apostle John writes to a young church surrounded by teachers who sound wonderful. “Beloved, do not believe every spirit, but test the spirits to see whether they are from God, for many false prophets have gone out into the world” (1 John 4:1, ESV). Notice the exact shape of the command, because it is the shape of everything this chapter taught. John does not say believe no one — that is not faith, it is the corrosion of it. He does not say believe everyone — that is not love, it is gullibility, and it gets the flock devoured. He says a third, harder thing: test. Weigh the claim against a fixed and trustworthy standard, and hold on to what proves true.
And why the warning? Because the false prophets were persuasive. That is what made them dangerous. Scripture is sober about this everywhere: some deceive, Paul says, “by smooth talk and flattery” (Romans 16:18, ESV) — by fluency, by the very polish that makes a thing easy to believe. The false teacher is not clumsy. He is articulate. Sound familiar? A large language model is fluent by design. Its confidence is manufactured, not earned; its polish is not evidence of truth — sometimes, as we saw in 8.3, it is precisely the disguise. Scripture named the danger of persuasive-and-wrong two thousand years before the transformer, and named the remedy too: test.
Testing requires a fixed reference, and this is the part the world forgets. You can only test a claim against something you trust more than the claim. For the Christian, that fixed point is the Word — this is what the Lutheran Confessions mean by sola Scriptura, Scripture as the norm that norms all other norms. The Bereans in Acts 17 were called noble — not rude, not faithless — precisely because they “examined the Scriptures daily to see if these things were so,” testing even the Apostle Paul’s preaching against the fixed standard. Testing the message was not an insult to the messenger; it was the highest honor they could pay the truth. Paul does not rebuke them. He commends them. The whole New Testament posture toward the confident claim is: “test everything; hold fast what is good” (1 Thessalonians 5:21, ESV).
Carry that straight into your work, and see how exactly it maps. Do not test one fluent voice against another fluent voice — do not fact-check the AI with the AI, do not confirm a rumor by finding it repeated. Test the claim against a fixed, trustworthy anchor: the primary source, the official record, the thing that is true whether or not anyone says it well. And notice the humility underneath John’s whole command. He assumes you can be deceived — “many false prophets have gone out.” The person most easily fooled is the one certain he cannot be. To test the spirits is to admit, up front, that a confident claim and a true claim are not the same thing, and that telling them apart is work — patient, unglamorous, accountable work. That admission is not weakness. In John’s letter and in your profession alike, it is the beginning of wisdom.
So: refuse to believe every spirit just because it is fluent. Test it. Hold fast to what is good. That is not cynicism and it is not credulity. It is the ancient, faithful third way — and it is, precisely, the discipline of a professional who owns the verdict.
8.11 — Common Pitfalls
Pitfall: Trusting an answer because it sounds authoritative. Example: A model returns a crisp statistic — “73% of small businesses adopted generative AI in year one” — in the same confident tone as everything else, and you put it on a leadership slide unverified. Fix: Tone is not evidence. Confidence and correctness are unrelated in an LLM (8.3). Any number that matters gets cross-checked against a trusted source before it leaves your hands.
Pitfall: Treating a citation as verified because it exists as text. Example: The AI cites “McKinsey Global Institute, The Small Business AI Index (2024).” It looks perfect, so you cite it too — and it was never published. Fix: A citation you did not open is a decoration, not a source (8.6). Search for it. Open it. Confirm it says what the AI claims it says.
Pitfall: Assuming the model’s knowledge is current. Example: You ask “what’s the best model right now?” and get a confident answer frozen at the model’s training cutoff, naming a tool that was superseded a year ago. Fix: For anything time-sensitive — prices, versions, laws, “latest” — assume the clock is stopped (8.4) and verify on a live, dated, trusted page yourself.
Pitfall: Fact-checking the AI with the AI. Example: You are unsure about a claim, so you ask the same tool “are you sure?” It says yes, confidently, and you feel reassured. Fix: A tool cannot be its own witness. Verify against an independent trusted source (8.7). “Are you sure?” often just triggers a more confident version of the same hallucination.
Pitfall: Letting the model talk you out of a correct answer. Example: You push back on something the AI got right; it apologizes and reverses to a wrong answer, because agreeable text is what it was trained to produce (sycophancy). Fix: If you can argue a model out of the truth, that tells you about the tool, not the truth. Anchor to a real source, not to which answer the model will defend under pressure.
Pitfall: Using AI in the red box — high stakes, hard to verify. Example: Someone asks a chatbot to confirm a medication dose, a legal deadline, or a safety limit and acts on the answer. Fix: Run the gate (8.8). High-stakes + hard-to-verify tasks do not go to AI; they go to a qualified human or the primary source. Know your red box and defend it.
Pitfall: Overcorrecting into total distrust. Example: After getting burned once, you stop using AI for anything and lose all its real value out of fear. Fix: The goal is calibration, not cynicism (8.9). Set the trust dial to the stakes: run free on low-stakes drafting, verify hard on high-stakes work. Refusing the tool is as unprofessional as blindly obeying it.
8.12 — Reps
The work is in the exercises. This is where verification stops being a thing you read about and becomes a thing your hands do. A preview of what is waiting:
- Break a model on purpose — ask it about a book, film, or law that does not exist and watch it invent one with total confidence.
- Dissect the sample answer — go through
code/ai-answer-with-errors.txtclaim by claim, classify each error, and prove the truth against a real source. - Run the five moves — take a real answer you need for work and put it through the verification toolkit from
code/verification-checklist.txt. - Test the training cutoff — find your tool’s frozen clock by asking about recent events, and catch a confidently-outdated “latest.”
- Build your own playbook — adapt the checklist into the one-page verification card you will use for the rest of the course.
A short Check Your Reps quiz is embedded on this page, right under the chapter — five questions drawn straight from what you just read. Take it before you move on. It is fair game for the midterm.
8.13 — This Week’s Project (Midterm)
This is the midterm week, and it comes in two parts. The graded midterm you sit is a separate, auto-graded quiz pool covering everything from Weeks 1–8 — history, models, prompting, tool choice, productivity, and this week’s limits. Study the whole first half of the book for it.
The take-home companion is P8 — “The Fact-Check Gauntlet,” specified in Project 8. It is a deliberately closed-AI challenge — because you cannot learn to verify AI by asking AI to do it for you. You will receive the fluent, error-laced briefing in code/ai-answer-with-errors.txt, and your job is to run the gauntlet: catch every seeded hallucination, classify it, verify the truth against a real source, produce a corrected version safe to ship, and — most important — write the verification playbook you will carry for the rest of the course. Normal tier catches and classifies the errors. Medium tier chases claims all the way to their primary sources. Hard tier is a judgment memo no AI can write for you: where should your team allow AI content to ship, and where must it never? That memo is where the thesis gets graded.
8.14 — Coach’s Final Word
Here is what I want you to carry out of Week 8. For seven weeks I taught you to trust this tool enough to use it. This week I taught you to distrust it enough to own it. Both are true at once, and holding both is the whole discipline. The professional who thrives is not the one who uses AI the most, and not the one who fears it the most. It is the one who has installed a reflex: fluent is a claim, not a fact — so test it.
Models will keep getting better at not lying. Good. But they will not stop lying, and even if they hallucinated half as often next year, the habit you build this week would still be the thing standing between you and the fabricated citation that ends up under your name. The tools turn over weekly. The discipline does not. Cross-check. Follow the source. Anchor to something you trust. And before anything ships — would I stake my name on this? If not, it does not go.
Underneath all of it runs the oldest instruction on the list. Do not believe every spirit, but test the spirits. Not because the world is only lies, and not because you are too clever to be fooled, but because a confident claim and a true one are different things, and telling them apart is the faithful, humble, accountable work you were made for. AI is an assistant, not an authority. You choose the tool. You own the verdict.
Now go do the reps. The Hunt is waiting right below this page, the sample answer is in code/, and the Gauntlet is where it all comes due.
See you on Monday.
Up next: Read the exercises and do all of Week 8’s reps, then run Project 8 — Project P8: The Fact-Check Gauntlet (the closed-AI midterm companion). Keep Appendix C open for the responsible-use rules behind the “red box,” lean on Appendix B for the research and answer engines that cite sources, and check any unfamiliar term against Appendix D. Then Chapter 9 — Creating with AI: Image and Video, where verification meets a new problem: telling what is real from what is generated.