The Model Landscape: Providers, Families, and Tradeoffs — Reps
Chapter 3 — Reps
Ten reps to turn “I use ChatGPT” into “I chose this tier for this task, and here’s why.” They move from mapping a lab’s family, to feeling the small-vs-large tradeoff in your own hands, to turning the five dials on your real work — the seed of this week’s project. No coding. Everything here happens in a browser, a phone app, or a spreadsheet. The keyboard is the gym. Don’t read these — run them.
Ground rules
- Type it yourself. Write your prompts and your notes by hand. Don’t paste this book’s words into a chatbot and call it a rep. Your fingers learn what your eyes skim.
- Run everything. Every prompt, every comparison, every dial. A rep you only imagined is a rep you didn’t do. Use the free tiers — they’re plenty for this (see Appendix A).
- Predict before you measure. Before you run a task or check a price, write down what you expect — which tier will win, how big the difference will be. Then compare. The gap between your guess and the result is the learning.
- AI usage note (every AI rep). End each rep that used a model with one honest line: what you asked, which provider + tier you used, and what you verified. You choose the tool; you own the verdict.
- Keep a
reps.txt. One short written reflection per rep, in a plain doc or notes file. This is graded thinking, not busywork. Never paste anything confidential into a consumer tool (see Appendix C).
Reps 1–3: Map the landscape
Rep 1 — Map one lab’s family
Pick one lab whose app you can open today (OpenAI/ChatGPT, Anthropic/Claude, or Google/Gemini). Open its model picker or its pricing/help page and write down its ladder: the cheapest/fastest option, the everyday default, and the top/flagship option. Where the app names a “reasoning,” “thinking,” “Pro,” or “deep research” mode, note it too. Keep code/providers-and-families.txt open as a reference, but fill this in from the live app.
Reflection: Was the ladder obvious in the app, or hidden? Most consumer apps nudge you toward one default and bury the choice. Which rung is the app’s default, and is that the rung you would pick for most of your work?
Rep 2 — Match five products to their engines
List five AI features you actually encounter — the “summarize” button in your email, a customer-service chatbot on a website, the writing helper in your word processor, a search tool, a note-taker. For each, guess (or look up) which foundation model or lab is under the hood and whether you’re even allowed to choose. It’s fine to be uncertain — mark your confidence.
Reflection: How many of your five let you see or choose the underlying model? What does it mean for your budget and your privacy when a product hides which engine it’s using and where your data goes?
Rep 3 — Price a task three ways (paper exercise)
Using the Claude ladder in code/providers-and-families.txt (Haiku ~$1→$5, Sonnet ~$3→$15, Opus ~$5→$25 per 1M tokens, mid-2026), estimate the rough relative cost of running one task — say, summarizing a one-page note — 10,000 times on each tier. You don’t need exact math; you need the ratio. Then flag every number you wrote as a “mid-2026 snapshot.”
Reflection: Output tokens cost several times input. Which is bigger for a summarize task — the input (the document) or the output (the short summary)? How does that change which tier you’d pick for high-volume summarizing?
Reps 4–6: Feel the tradeoff
Rep 4 — Small vs flagship on a simple task
Pick a genuinely simple task: reword a two-line email politely, or turn five bullet points into a paragraph. Run it on a small/fast tier and on a flagship tier (or on two different assistants where one is clearly the “lighter” option). Judge honestly: can you tell which answer came from the bigger model without looking?
Reflection: Could you tell the difference? If not, what did paying for the flagship actually buy you on this task? End with your AI usage note.
Rep 5 — Small vs flagship on a hard task
Now pick a genuinely hard task with real steps: analyze the pros and cons of a decision with three competing factors, or work through a multi-part planning problem from your job. Run it on a small tier and a flagship (or a standard mode vs a reasoning mode). This time, look for where the weaker one breaks — a dropped constraint, a shallow answer, a logical slip.
Reflection: Where exactly did the smaller/standard model fall short — and would you have caught that slip if you weren’t comparing? Tie this to the spine rule: you own the verdict on both answers. End with your AI usage note.
Rep 6 — Toggle reasoning on and off
Take one shallow task (format a list) and one deep task (a multi-step logic or analysis problem). Run each with a reasoning/thinking/Pro mode on, then off. Note the difference in time and in quality for each.
Reflection: On which task did reasoning mode earn its extra time and cost — and on which was it pure waste? State the rule you’d give a coworker in one sentence. End with your AI usage note.
Reps 7–9: Turn the five dials
Rep 7 — Score three real tasks on the five dials
Pick three real tasks from your own week. For each, rate the five dials from Chapter 3.5: difficulty (1–5), speed need (low/med/high), budget sensitivity (low/med/high), privacy sensitivity (low/med/high), and multimodal? (does it involve images/audio/video?). Predict the recommended tier before you decide.
Task: _______________________________________________
Difficulty (1-5): ___ Speed: ___ Budget: ___
Privacy: ___ Multimodal? ___ -> My tier guess: _______
Reflection: Which dial did the most work in your decision for each task? For most everyday professional work, is your honest answer higher or lower on the ladder than your instinct wanted?
Rep 8 — Run the Model Chooser and argue with it
Open the Model Chooser widget below Chapter 3. Enter each of your three tasks from Rep 7 and compare its recommendation to your prediction. Then change one dial per task and watch the recommendation move.
Reflection: Where did you and the widget disagree — and who was right, once you thought it through? (The widget is a helper, not an authority; you own the verdict.) Which single dial change flipped a recommendation most dramatically?
Rep 9 — Find the task where privacy overrides everything
Identify one real task involving information you should not paste into a public consumer chatbot (a client’s data, a patient detail, an unreleased plan — skim Appendix C). Set its privacy dial to high and note how that overrides the other dials — even if a cloud flagship is smarter, privacy sends you to a governed enterprise tool or a local model (previewed for Chapter 11).
Reflection: Describe the task in a way that reveals nothing confidential. Why is choosing a “less capable” private option here the wiser choice, not the weaker one? Tie it to counting the cost (Chapter 3.8).
Rep 10: Seed the project
Rep 10 — Start your model-selection matrix
Open code/model-selection-matrix.csv in a spreadsheet app (Google Sheets, Excel, or Numbers — free options in Appendix A). Study the EXAMPLE row, then fill in two of the six task rows with real tasks from your job, turning all five dials and writing a one- or two-sentence justification for each on cost/speed/quality.
Reflection: Which of your two tasks was hardest to place, and why? Placing the ambiguous task well — the one that could go two ways — is the real skill this week is building.
Done? One Last Thing.
A miniature of Project 3, end to end. Take one real recurring task from your job. In a single page (a doc or a slide):
- Describe the task in one sentence — no confidential details.
- Turn all five dials and write the numbers/levels down.
- Name a provider and a tier you’d reach for — pin the model ID where you can (e.g.,
claude-sonnet-5), and date it “mid-2026.” - Justify it in three lines — one each for cost, speed, and quality — plus a word on the privacy call.
- State the cost you counted: why not the flagship (or why it is worth it here).
If you can do this cleanly for one task tonight, you can do it for six in Project 3. That’s the whole job in miniature: turn the dials, choose the rung, count the cost, own the verdict.
Up next: Project 3 — Project P3: The Model-Selection Matrix.