Chapter 03 · Week 3

The Model Landscape: Providers, Families, and Tradeoffs

What does it mean to count the cost — to value a tool rightly instead of always reaching for the most expensive one?

Chapter 3 — The Model Landscape: Providers, Families, and Tradeoffs

“If all you have is a hammer, everything looks like a nail.” — the law of the instrument

“For which of you, desiring to build a tower, does not first sit down and count the cost, whether he has enough to complete it?” — Luke 14:28 (ESV)


Why This Matters

Two weeks ago you learned what a language model is. Last week you fed one prompt to three different assistants and watched them disagree. This week we zoom out to the whole board, because there is a quiet fact that most working professionals never learn, and it costs them money, time, and quality every single day: there is not “an AI.” There are dozens — made by a handful of labs, each shipping a whole family of models — and reaching for the same one every time is like owning a full toolbox and using the hammer for everything.

Here is how most people use AI in 2026. They open ChatGPT. They type. They close it. “I use ChatGPT” has become a sentence the way “I’ll Google it” once was — one brand standing in for an entire category. And that habit is not wrong, exactly; ChatGPT is a fine default. But it is a habit of a person who does not yet know the landscape, and it quietly leaks value. It reaches for the most expensive engine to answer a one-line question. It uses a text-first tool for a job a multimodal one would crush. It pastes a client’s confidential document into a consumer chatbot because “the AI” was the only tool the person knew how to hold.

The professional who thrives with AI is not the one who uses it the most. It is the one who knows which tool to reach for — and that is a real, learnable skill. It has a shape. Every major lab builds a ladder of models, from a tiny, fast, dirt-cheap one at the bottom to a giant, slow, expensive, brilliant one at the top. The whole game of Chapter 3 is teaching your hand to reach for the right rung.

There are two sides to this, and both are your job now. The first is the tool you wield: choosing well makes your work faster, cheaper, and better. The second is the cost you govern: when you have a company AI budget, or when your employer is watching what you paste and where it goes, the model you pick has a bill and a risk attached. Non-technical professionals are increasingly the ones holding that budget line and making that call. Choosing badly is now a way to waste real money and expose real secrets. Choosing well is a competency people will pay you for.

And underneath all of it runs this week’s question — the one Luke puts in the mouth of a man about to build a tower. What does it mean to count the cost — to value a tool rightly, instead of always reaching for the biggest, most expensive one? Hold that. We will earn it near the end.

Coach’s Note — Say the spine rule with me, because this is the week it gets literal: You choose the tool. You own the verdict. AI drafts, generates, and accelerates — but the choice of which AI, and the judgment about whether its answer is any good, stays with you. This chapter is entirely about the first half of that sentence. Next chapter is about the second. Neither is ever the machine’s to make.


3.1 — Foundation Models: The Engines Behind the Apps

Start with a distinction that clears up half the confusion in this field. There is a difference between the app you open and the model that answers you.

The app is the friendly front door: ChatGPT, the Claude app, the Gemini app, Copilot in your Office ribbon, Perplexity’s search box. It has a login, a chat window, buttons, a monthly price. The model is the engine humming behind that door — the trained system, built at enormous expense, that actually reads your words and produces the reply. One is the car’s dashboard; the other is the engine under the hood.

Those engines have a name: foundation models. A foundation model is a large, general-purpose model trained once, at great cost, on a vast amount of data, and then reused as the base for countless tasks and products. This is the important part for you as a professional: a handful of labs build these engines, and then thousands of apps — including apps that don’t say “AI” anywhere on the label — put those same engines under their own hoods. The customer-service chatbot on a retailer’s website, the “summarize this thread” button in your email, the writing helper in your word processor: many of them are one of five or six foundation models wearing a different coat.

Why does this matter to someone who will never train a model? Because it means the landscape is smaller and more learnable than it looks. You do not have to keep up with ten thousand AI products. You have to know the handful of labs that make the engines, and the shape of the family each one ships. Learn those, and every new product that crosses your desk snaps into a picture you already understand: “Oh — that’s a mid-tier Google model in a nice wrapper.”

Coach’s Note — When a vendor pitches you their shiny new “AI-powered” tool, one of your first questions should be: which foundation model is under the hood, and can I choose? Sometimes the magic is a genuinely clever product. Just as often it is gpt-5.6 or claude-sonnet-5 with a logo on it and a markup on the price. Knowing the engines lets you tell the difference.


3.2 — Meet the Frontier Labs

A frontier lab is one of the small number of organizations pushing the leading edge of what these models can do. As of mid-2026, this is roughly the field. (Every fact in this section moves fast — dates, versions, and prices are mid-2026 snapshots. Learn the reputations; they age better than the version numbers.)

LabModel familyPowersKnown for (mid-2026)
OpenAIGPT-5.xChatGPT, CopilotThe default all-rounder; largest reach. Strong writing, images, voice, agentic “computer use.”
AnthropicClaudeClaude apps, Claude CodeCareful reasoning, long documents, top coding; a “trust and precision” reputation; ~1M-token memory.
GoogleGemini 3.xGemini, WorkspaceBest native multimodal (reads video + audio directly); deep Gmail/Docs/Sheets integration; ~1M context.
xAIGrok 4.6Grok (in X)Real-time info from the live web; cheaper tokens; looser content rules.
MetaLlama 4 (Scout/Maverick)self-hosting, many appsOpen-weight — free to download and run yourself.
MistralMistral Large 3its apps + self-hostingEuropean; open-weight-friendly.
DeepSeekDeepSeek V4its app + self-hostingNear-frontier quality at a fraction of the cost; Chinese lab (raise data-residency questions).
AlibabaQwen3self-hosting, its appsStrong open-weight models and embeddings.
MicrosoftPhi (small), MAICopilot, AzureSmall, efficient, on-device-friendly models; also resells the others through Azure.

A few things to notice, because they are the durable lessons and not the disposable names:

There is no single “best” in 2026. This is the sentence I most want you to carry out of this chapter. The field has specialized. As a rough map: ChatGPT leads on reach and being an all-rounder; Claude leads on careful reasoning, long documents, and coding; Gemini leads on multimodal and Google-Workspace integration; Grok leads on real-time information; the open-weight labs (Meta, Mistral, DeepSeek, Alibaba) lead on privacy and cost because you can run them yourself. Anyone who tells you one brand is simply “the best” is selling something or hasn’t looked lately.

Some engines you rent; some you can own. OpenAI, Anthropic, and Google keep their model weights private — you rent access through an app or an interface. Meta, Mistral, DeepSeek, Alibaba, and Microsoft’s Phi line release their weights so you can download and run them yourself. That single difference — closed versus open-weight — drives cost and privacy, and we’ll pull on it in 3.7.

A word of care on the fast-moving names. I have pinned specific versions here so the picture is concrete, but a validator, a vendor page, and next Tuesday will all disagree with some of them. Two examples the record should be honest about: Meta previewed a giant “Llama 4 Behemoth” that, as of this writing, never actually shipped — so don’t cite it as a tool you can use. And prices below are snapshots; the vendor console is the truth. When in doubt, teach yourself the category, not the SKU.


3.3 — Why Every Lab Ships a Family, Not a Model

Open any lab’s menu and you will not find one model. You will find a lineup — a ladder of options with names like Nano, Mini, Flash, Sonnet, Opus, Pro. This confuses people. It shouldn’t. It is the most useful structure in the whole landscape, and it is exactly like things you already buy.

Think of shipping a package. You don’t have “one way to ship.” You have overnight, two-day, and ground. Overnight is fast and expensive; ground is slow and cheap. You choose per package based on how much the speed is worth. Or think of travel: economy, business, first. Same plane, same destination — you pay for the tier the trip is worth. A model family is that menu. The labs build a range from cheap-fast-simple to expensive-slow-powerful so you can match the model to the worth of the task.

Here’s the recurring shape of the ladder, with real examples as of mid-2026:

RungWhat it’s forClaude exampleOthers (mid-2026)
Nano / TinyCheapest, fastest; huge volume of simple work(Anthropic markets fewer public tiers)gpt-5.6-nano
Mini / SmallCost-efficient; most everyday tasksClaude Haiku 4.5 (claude-haiku-4-5)gpt-5.6-mini, gemini-3-flash
BalancedThe sane default for most workClaude Sonnet 5 (claude-sonnet-5)gpt-5.6, gemini-3-pro
FlagshipHard, multi-step work; long documents; agentsClaude Opus 5 (claude-opus-5)gpt-5.6
Reasoning / Pro”Thinks” before answering; deep analysisClaude Opus 5 in extended thinkinggpt-5.6-pro, gemini-3-pro
Frontier / CreativeThe top endClaude Fable 5 (claude-fable-5)
Realtime / VoiceLive spoken conversation, low latency(voice modes in the app)ChatGPT Advanced Voice

The specific names will rotate. The ladder will not. Every capable lab has a cheap-fast bottom, a balanced middle, and an expensive-powerful top, plus a few specialists (a reasoning mode, a voice mode). When you meet a new lab, the first thing to do is find its ladder — ask “what’s the fast/cheap one, what’s the default, what’s the top?” — and you have oriented yourself in thirty seconds.

Here is the Claude ladder with the known-stable IDs, because you’ll use it as your reference all term. (Prices are per one million tokens, a “token” being roughly three-quarters of a word — see Appendix D. These are mid-2026 snapshots and output costs several times input.)

Model (pin this ID)Rough $ per 1M in → outReach for it when…
Claude Haiku 4.5 (claude-haiku-4-5)~$1 → ~$5high-volume, simple: classify, extract, short replies
Claude Sonnet 5 (claude-sonnet-5)~$3 → ~$15the everyday default for most professional work
Claude Opus 5 (claude-opus-5)~$5 → ~$25hard reasoning, long documents, multi-step tasks
Claude Fable 5 (claude-fable-5)~$10 → ~$50frontier/creative top end — the newest rung (mid-2026 snapshot)

The ladder reads bottom to top as cheap-and-simple → expensive-and-powerful. Your skill is knowing which rung a given task actually needs — and, most of the time, the honest answer is lower than your instinct wants to reach.


3.4 — Why Smaller Models Exist (and Why They’re Often the Right Answer)

If the flagship is the smartest, why does anyone use the small ones? New professionals assume the small models are just the “budget version for people who can’t afford the good one.” That is exactly backwards, and correcting it is one of the most valuable things this chapter can do for you.

Small models exist because most work is not hard. Reply to a routine email. Classify a support ticket as “billing” or “technical.” Pull the three action items out of a meeting note. Turn a bulleted list into a tidy paragraph. None of that needs a giant reasoning engine any more than mailing a birthday card needs an overnight courier. A small model does these jobs faster, cheaper, and — this is the part people miss — often just as well, because the task had no depth for the big model’s extra brain to show off in.

Smaller models win on four things at once:

  • Cost. A flagship can cost 10 to 100 times more per token than a nano/mini model. For high-volume work, that’s the difference between a rounding error and a shocking invoice.
  • Speed and latency. Smaller models answer faster and with less lag. For anything interactive — a live chat, a voice assistant, a “make this snappy” tool — speed is the quality.
  • “Good enough” is a real standard. For a huge share of daily tasks, a mini model’s answer is indistinguishable from a flagship’s. Paying more buys you nothing you can see.
  • They can run on small hardware. The smallest open-weight models fit on a laptop or a phone — which is exactly what makes the private, offline, no-per-token-cost path in Chapter 11 possible.

The flip side is honest too. Small models genuinely are weaker at hard, multi-step reasoning, long documents, subtle judgment, and staying coherent over a very long task. The skill is not “always go cheap” or “always go big.” It is matching. Send the postcard by ground and the contract by courier.

Coach’s Note — The most expensive habit in professional AI use in 2026 is reflexively reaching for the flagship. It feels responsible — “I used the best one!” — but for trivial work it is the opposite of responsible: you paid frontier price and waited longer for an answer a mini model would have nailed in half the time. We will name this exact sin in the apologetic section. It has an old name: failing to count the cost.


3.5 — The Five Dials: Cost, Speed, Reasoning, Memory, Multimodal

When you choose a model, you are really turning five dials. Learn the dials and you can reason about any model, including ones invented after this book is printed.

1. Cost. How much does each use cost, and how many times will you do it? A pricier model on a one-off task is cheap in absolute terms. A pricier model run ten thousand times a day is a budget crisis. Two facts to hold: flagships cost 10–100× a small model per token, and output tokens dominate the bill — the model’s reply typically costs four to eight times as much as your prompt. A tool that writes long, rambling answers is quietly expensive.

2. Speed / latency. How fast does the first word appear, and how fast does it finish? Bigger models are slower. For batch work you run overnight, you don’t care. For a live conversation or a customer-facing chat, a two-second lag is the whole user experience. Speed is not a luxury; for interactive work it is the product.

3. Reasoning ability. How hard is the thinking? A model’s depth shows up only on tasks that have depth — multi-step logic, careful analysis, a tricky legal or financial nuance, a long chain of “if this then that.” On a shallow task, a flagship’s extra reasoning is horsepower with nowhere to go.

4. Memory (context window). How much can the model hold in its head at once? This is the context window, measured in tokens (Chapter 2). As of mid-2026, flagship Claude and Gemini reach roughly one million tokens — enough to hold a long contract, an entire report, or several books in a single session. That’s transformative for “read this whole thing and answer questions about it” work. But big context is slower and pricier, so you don’t pay for a million-token model to summarize a paragraph.

5. Multimodal. What kinds of stuff can it handle — just text, or also images, audio, and video, in and out? “Multimodal” means more than one mode. If your task involves a photo, a chart, a screenshot, a voice clip, or a video, you need a model built for that mode. As of mid-2026, Gemini is the standout for reading video and audio natively; image and video generation is a specialist world we visit in Chapters 9 and 10.

You will practice turning these dials by hand in this week’s project, and the Interactive Lab lets you turn them and watch the recommendation change in real time. Downstream — in the widget, the reps, and the project — we operationalize these as task difficulty (reasoning + memory folded together), speed, budget (cost), and multimodal, plus one more dial that outranks them all when it fires: privacy (3.7).

Coach’s Note — Notice that four of the five dials often point down the ladder, not up. Cheaper, faster, and small-hardware-friendly all favor the smaller model; only “hard reasoning” and “huge memory” push you toward the flagship. That’s why the honest default for everyday work sits in the middle of the ladder, not the top. Start at balanced, drop to mini when the task is simple, climb to flagship only when the task earns it.


3.6 — Reasoning Models vs Standard Models

One split on the ladder deserves its own section because it is new, powerful, and easy to misuse: the difference between a standard model and a reasoning model.

A standard model answers the way you’d answer a question you already know — it responds more or less immediately. A reasoning model is trained to think first: before it gives you an answer, it works through the problem in a series of internal steps, checking itself along the way, the way you’d scribble on scratch paper before committing to a hard answer. As of mid-2026 you’ll see this as a “reasoning” or “thinking” or “Pro” or “deep research” mode — gpt-5.6-pro, gemini-3-pro, Claude Opus in extended thinking. Some apps let you toggle it on; some route to it automatically for hard questions.

When is the extra thinking worth it? When the task genuinely has steps: a multi-part logic puzzle, a careful financial or legal analysis, planning a complex project, a research question that requires chasing and weighing several sources, math that must actually be right. On those, a reasoning model is dramatically better — it catches its own mistakes mid-thought.

When is it a waste? On everything shallow, which is most things. Reasoning mode is slower (it’s literally doing more work) and more expensive (all that internal thinking is tokens you pay for — a single question can trigger many model calls behind the scenes). Turning on deep-reasoning mode to reword an email is like hiring a forensic accountant to split a dinner check. It’ll get the right answer. It’ll also take longer and cost more, for a job that never needed it.

Here’s the mental rule: standard for recall and routine; reasoning for problems with steps. If you could plausibly do the task in one pass yourself, a standard model is fine. If you’d need scratch paper, reach for reasoning.

Coach’s Note — Reasoning models are also where the “deep research” agents live — the ones that browse dozens of sources and hand you a cited report. They are genuinely impressive and genuinely fallible: they can chase a wrong thread confidently for pages, and they can invent citations that look perfect. The extra thinking makes them stronger, not infallible. Everything from Chapter 2 still holds — verify before you trust. We drill exactly this in Chapter 7 and Chapter 8.


3.7 — Open vs Closed: A Preview of Where Your Data Lives

Now the split that will matter most the day you have something confidential to process. Models come in two broad flavors, and the difference is about who holds the engine.

Closed / commercial models — GPT, Claude, Gemini — keep their weights private. You rent access: you send your words to the lab’s computers, the model runs there, the answer comes back. These are usually the newest, largest, and smartest models available. The tradeoff is that your text left your building. For public or low-sensitivity work, that’s fine. For a client’s confidential file, it is a question you must answer before you paste (Chapter 11 and Appendix C).

Open-weight models — Llama 4, Mistral, DeepSeek, Qwen, Microsoft’s Phi — publish their trained weights. You can download them and run them on your own computer or your company’s servers. That means the data never leaves your machine (privacy, offline, no per-token fee). The tradeoff: the open-weight models you can realistically run on normal hardware are usually a step behind the top closed flagships, and you need a decent machine to run the good ones.

Three points of care, because this is a place people say sloppy things:

  • “Open-weight” is not “open source.” You get the finished, trained model — rarely the training data or the full recipe. It’s a downloadable engine, not a fully open blueprint.
  • Licenses differ and they matter. Some open-weight models ship under permissive licenses (Apache-2.0, MIT — e.g., much of Mistral and DeepSeek); others use model-specific “community” licenses with conditions. If you’ll use one commercially, someone needs to actually read the license.
  • Where the lab sits is part of the calculation. A capable, low-cost model from an overseas lab (DeepSeek, Qwen) is a real option — and raises a legitimate data-residency question for a professional handling regulated or sensitive information. That’s not xenophobia; it’s the same due diligence you’d do for any vendor.

This is a preview. You will actually run a local model with your own hands in Chapter 11, and weigh the privacy-versus-capability tradeoff for real. For now, plant one idea: sometimes the “worse” model is the right choice, because it keeps the data home. Capability is not the only axis that matters.

Coach’s Note — File this next to the cost lesson: the biggest, smartest, newest model is not automatically the right one. Sometimes a smaller model is right because it’s cheaper. Sometimes an open-weight model is right because it’s private. “Best on the benchmark” and “best for this task, this budget, and this data” are different questions, and only the second one is yours to answer.


3.x — Interactive Lab: Model Chooser

Below this chapter on the website you’ll find an interactive panel called the Model Chooser. Go use it now — it’s not decoration, it’s the rep that wires today’s lesson into your hands.

The Chooser gives you the five dials from 3.5 as controls: a difficulty slider, a speed need, a budget sensitivity, and a privacy sensitivity (plus whether the task is high-volume). You set them to describe a real task, and the panel recommends a tier — nano/mini, balanced, flagship/reasoning, or “keep it local” — and, crucially, shows you the tradeoff you’re accepting: the cost you’d save or spend, the speed you’d gain or lose, the capability you’d give up.

What it teaches is not a lookup table of “task X → model Y.” It teaches the feel of the tradeoff — how turning up “difficulty” pushes you up the ladder while turning up “budget sensitivity” and “volume” pushes you back down, and how “privacy” can override everything and send you local even when the cloud model is smarter. Play with it deliberately. Set a task, predict the recommendation before you release the last dial, then see if you were right. Then change one dial and watch the recommendation move. That reflex — “which dial did I just turn, and which way did it push me?” — is the entire skill of model selection, and it’s exactly what you’ll perform by hand in this week’s project.

In BoodleBox. The Chooser teaches the feel of the dials; BoodleBox shows you the actual menu you’re choosing from. In a chat, open the bot picker and select the “AI Models” category to see the real engines Concordia has put on your desk — as of 2026, models like ChatGPT 5.6, Claude Opus 5, Gemini 3.1 Pro, Perplexity, and the image models — the concrete face of the ladder you just turned. Use the Chooser to settle on a rung, then @-mention the matching model for that task, and switch models per task instead of firing everything at one default. And because BoodleBox leans on heavy token-reduction to keep costs down (as of 2026), reaching for the cheaper rung is even easier there than the raw price ladder implies — the affordability dial is quietly turned in your favor. Fallback: no BoodleBox on hand? The same choose-then-run reflex works in any single assistant’s own model picker.


3.8 — Counting the Cost

Now the week’s question, given its due. What does it mean to count the cost — to value a tool rightly, instead of always reaching for the biggest, most expensive one?

Jesus asks it in Luke 14, and He is not talking about model pricing; let’s be honest about the text before we borrow from it. He is teaching about the cost of following Him. “For which of you, desiring to build a tower, does not first sit down and count the cost, whether he has enough to complete it?” (Luke 14:28, ESV). He follows it with a king who sits down to reckon whether ten thousand can meet twenty thousand before he marches. The point in context is sober and severe: discipleship is not a thing you drift into on a wave of feeling. You sit down. You reckon. You count what it will require and whether you will finish. The half-built tower — the abandoned project, the enthusiasm that could not pay its own bill — is a picture of a faith begun without weighing.

So the first thing to say is that counting the cost is not a budgeting tip with a verse taped to it. It is a posture Jesus commends: deliberateness before commitment, honesty about what a thing will require, the refusal to begin what you have not weighed. That posture is a creational good — a piece of the wisdom God built into ordering our lives — and because it is that kind of good, it reaches into the small, ordinary places too, including how a Christian professional spends money, attention, energy, and trust on a Tuesday afternoon.

And here is where it presses on this chapter, precisely. The reflex to always reach for the biggest, newest, most expensive model is a failure to sit down and count. It feels responsible — I used the best one — but it is thoughtlessness wearing the costume of diligence. You spent your employer’s money you didn’t need to spend. You burned more of the shared thing — the electricity, the water that cools the datacenter, the strained grid — than the task required, because more felt safer than considered. The frontier model idles its enormous brain over a task a mini model would have finished, and you call the waste “quality.” Scripture has an unfashionable word for the opposite of this waste, and it is not stinginess. It is stewardship — faithful management of what belongs to another. The tower still gets built in Jesus’ story; the wise builder is not the one who refuses to build but the one who counts first and finishes.

Notice, too, what counting the cost frees you from. Our moment has an idolatry of more — bigger, newer, maxed-out, top-tier — as if reaching for the largest thing were itself a virtue. The tower-builder’s wisdom quietly dethrones that. The right tool is not the most impressive tool; it is the one fitted to the task and the trust you carry. To choose a humble model on purpose — because the job is small, or because the data must stay home — is not settling. It is discernment. It is the same wisdom that keeps a good manager from buying a forklift to move a single box.

Which loops us back to the spine rule, and shows it was never merely a productivity slogan. You choose the tool. You own the verdict. The choosing is a moral act, not only a technical one, because you are accountable for what you spend, what you expose, and what you waste in the choosing — the neighbor’s budget, the neighbor’s confidential file, the creation’s finite resources. In the old Lutheran language, this ordinary desk work is vocation: a station in which God has placed you to serve your neighbor faithfully. The neighbor is served not by your reaching for the flashiest engine but by your counting the cost and choosing well. Sit down first. Reckon. Then build.


3.9 — Common Pitfalls

Pitfall: Treating one brand as “the AI” and reaching for it every time. Example: You paste every task — a birthday note, a legal summary, a photo to describe — into the same chatbot, because it’s the only door you know how to open. Fix: Learn the labs and the ladder (3.2, 3.3). Keep the code/providers-and-families.txt sheet handy and ask, per task, which engine and which rung this actually needs.


Pitfall: Always reaching for the flagship because “the best one” feels responsible. Example: You run a top-tier reasoning model to reword a two-line email, thousands of times a month, then flinch at the bill and the lag. Fix: Start at the balanced rung; drop to mini for simple, high-volume work; climb to flagship only when the task has real depth. Match the tool to the worth of the job (3.4).


Pitfall: Turning on deep-reasoning mode for shallow tasks. Example: You leave “extended thinking / Pro” on for everything, so a one-line formatting request takes fifteen seconds and several times the cost. Fix: Standard for recall and routine; reasoning only for problems with real steps (3.6). If you wouldn’t need scratch paper, the model doesn’t either.


Pitfall: Ignoring the privacy dial and pasting sensitive data into whatever’s open. Example: You drop a client’s confidential contract into a consumer chatbot because “the AI” was the tool at hand, and the text has now left your building. Fix: Treat privacy as a dial that can override capability (3.5, 3.7). For sensitive material, use a governed enterprise tool or a local model (Chapter 11), and read Appendix C first.


Pitfall: Trusting a fast-moving number as if it were permanent. Example: You quote a model’s price or a version from a six-month-old article and build a budget or a decision on it, not knowing the price dropped and the model was replaced. Fix: Treat every price, version, and “best” as a dated snapshot. Pin model IDs (claude-sonnet-5), not marketing names, and confirm on the vendor’s own pricing page before it matters.


Pitfall: Confusing “open-weight” with “free of obligations” or with “open source.” Example: You build a commercial product on a downloaded model without reading its license, assuming “open” means “do anything.” Fix: Open-weight means you get the trained engine, not a blank check. Check the license (Apache-2.0 and MIT are permissive; community licenses have conditions) and the data-residency question before you commit (3.7).


Pitfall: Assuming the biggest model is automatically the most accurate. Example: You accept a flagship’s confident answer without checking because it’s the “smartest” model, and it was confidently wrong — as any tier can be. Fix: Capability reduces error; it does not remove it. Every tier hallucinates. You own the verdict regardless of which rung produced the draft (spine rule; drilled hard in Chapter 8).


3.10 — Reps

The work is in the exercises. The keyboard is the gym; this is where Week 3 gets into your hands. You’ll do these in BoodleBox (or a public assistant as fallback) and a spreadsheet — no coding. A preview of what’s waiting:

  • Map the ladder for one lab you actually use — find its cheap/default/top rungs and write them down.
  • Run the same simple task on a small tier and a flagship and judge honestly whether you can tell the difference.
  • Turn each of the five dials on three real tasks from your own week and predict the tier before you check.
  • Toggle a reasoning mode on and off for one shallow and one deep task, and time both.
  • Start your model-selection matrix — the seed of this week’s project — using code/model-selection-matrix.csv.

Each rep ends with a short written Reflection, and every rep that uses AI ends with an honest one-line AI usage note: what you asked, which tier you used, and what you verified. A short Check Your Reps quiz is embedded on this page, right under the chapter — five questions grounded in exactly what you just read. Take it before you move on.


3.11 — This Week’s Project

Your project is P3 — “The Model-Selection Matrix,” specified in Project 3. You’ll take six real tasks from your own job, run each through the five dials, and choose a provider and tier for each — justifying every choice on cost, speed, and quality, and noting the privacy call. You’ll fill the provided matrix (code/model-selection-matrix.csv) using the one-page reference (code/providers-and-families.txt) as your guide.

At a high level: Normal tier fills the matrix for six tasks with defensible justifications. Medium tier stress-tests one choice by actually running the task on two different tiers and comparing. Hard tier is a one-page memo to the Rivertown office manager (the project’s scenario) recommending a sensible default model policy — which tier to reach for first, and when to climb or drop — and defending it. That memo is the judgment no model can produce for you, and it’s where the thesis gets graded.


3.12 — Coach’s Final Word

Here’s what I want you to carry out of Week 3. The person who wins with AI is not the one with the most subscriptions or the habit of always clicking the “best” model. It’s the one who can look at a task, turn five dials in their head — how hard, how fast, how cheap, how private, what kind of stuff — and reach, without drama, for the right rung on the right ladder. That’s it. That’s the skill. It looks small and it saves fortunes.

The names in this chapter will rot. Half the version numbers will be higher by the time you read this twice, and a lab or two will have leapfrogged another. Let them. You are not memorizing a leaderboard. You are learning a shape — labs that make engines, families that ladder from cheap-fast to expensive-powerful, five dials that decide the rung — and that shape will still be true when every SKU in here is history.

And underneath the skill runs the tower-builder’s wisdom, which is older than every model and will outlast them all. Sit down first. Count the cost. Don’t reach for the biggest, most expensive tool because bigger feels safer; reach for the one fitted to the task and the trust you carry — the employer’s money, the neighbor’s confidence, the creation you’re spending to get the answer. To choose a humble tool on purpose is not settling; it’s stewardship. The tower still gets built. The wise builder is just the one who counted first.

Now go do the reps. The Model Chooser is waiting right below this page, the reference sheet is in code/, and Project 3 is where it all comes together.

See you on Monday.


Up next: Read the exercises and do all of Week 3’s reps, then build Project 3 — Project P3: The Model-Selection Matrix. Sign in to BoodleBox per Appendix A (it also covers the free fallbacks), keep the tool directory in Appendix B open as you choose, review the data-sensitivity rules in Appendix C before you paste anything, and check any unfamiliar term against the Appendix D glossary. Then Chapter 4 — Prompting as Professional Communication.

Interactive Lab — Week 3
Model Chooser

Describe a real task by turning the five dials from Section 3.5. The panel picks a rung on the ladder and shows the tradeoff you're accepting. Predict the recommendation before you move the last dial — then change one dial and watch it move. That reflex is the whole skill.

Everyday
ShallowDeep steps
Some
Don't careEvery cent
Normal
Can waitInstant
Internal
PublicConfidential
Recommended rung
Balanced
the sane default for most work

Relative cost
Relative speed
Try: Set difficulty to Deep steps and watch it climb to a reasoning model. Now drag Budget sensitivity and High volume up — see it pushed back down the ladder. Then set Privacy to Confidential: capability stops mattering and it sends you local.
Check Your Reps

Check Your Reps — The Model Landscape

Question 1 of 5
The chapter separates the "app" you open from the "model" that answers you. Which best describes a foundation model?
Why: A foundation model is the trained engine built once and reused under the hood of many different apps, which is why a handful of labs' engines power thousands of products.
Question 2 of 5
According to the chapter, what is the most expensive habit in professional AI use in 2026?
Why: Reaching for the flagship on shallow work feels responsible but pays frontier prices and waits longer for an answer a mini model would have nailed — a failure to count the cost.
Question 3 of 5
The chapter notes that "output tokens dominate the bill." What is the practical implication?
Why: Because the model's reply typically costs several times more than your prompt, verbose answers run up the bill even when the question was short.
Question 4 of 5
When is choosing a reasoning ("thinking") model worth it over a standard model?
Why: Reasoning models shine on tasks that actually have steps you'd need scratch paper for; on shallow work they're just slower and more expensive for no visible gain.
Question 5 of 5
A colleague must process a client's confidential contract. Based on the chapter's "privacy dial," what should most guide the choice?
Why: Privacy is a dial that can override capability: for confidential data, keeping it home with a governed or local model matters more than reaching for the smartest cloud model.
YOU FINISHED. NICE WORK.