Chapter 03 · Reps

The Model Landscape: Providers, Families, and Tradeoffs — Reps

← Back to Chapter 3

Chapter 3 — Reps

Ten reps to turn “I use ChatGPT” into “I chose this tier for this task, and here’s why.” They move from mapping a lab’s family, to feeling the small-vs-large tradeoff in your own hands, to turning the five dials on your real work — the seed of this week’s project. No coding. Everything here happens in a browser, a phone app, or a spreadsheet. The keyboard is the gym. Don’t read these — run them.

Ground rules

  • Type it yourself. Write your prompts and your notes by hand. Don’t paste this book’s words into a chatbot and call it a rep. Your fingers learn what your eyes skim.
  • Run everything. Every prompt, every comparison, every dial. A rep you only imagined is a rep you didn’t do. Run them in BoodleBox (sign in with your Concordia account — setup in Appendix A); its multi-bot chats let you put several models in one conversation, which is exactly what this week’s comparison reps need. Fallback: the free tiers of the public assistants work too.
  • Predict before you measure. Before you run a task or check a price, write down what you expect — which tier will win, how big the difference will be. Then compare. The gap between your guess and the result is the learning.
  • AI usage note (every AI rep). End each rep that used a model with one honest line: what you asked, which provider + tier you used, and what you verified. You choose the tool; you own the verdict.
  • Keep a reps.txt. One short written reflection per rep, in a plain doc or notes file. This is graded thinking, not busywork. Never paste anything confidential into a consumer tool (see Appendix C).

Reps 1–3: Map the landscape

Rep 1 — Map one lab’s family

Pick one lab whose app you can open today (OpenAI/ChatGPT, Anthropic/Claude, or Google/Gemini). Open its model picker or its pricing/help page and write down its ladder: the cheapest/fastest option, the everyday default, and the top/flagship option. Where the app names a “reasoning,” “thinking,” “Pro,” or “deep research” mode, note it too. Keep code/providers-and-families.txt open as a reference, but fill this in from the live app.

In BoodleBox. Open the bot picker and choose the “AI Models” category to see, in one place, the engines Concordia has licensed across labs — ChatGPT, Claude, Gemini, Perplexity, image models (exact versions move; as of 2026 you’ll see the likes of ChatGPT 5.6 and Claude Opus 5). That’s the cross-lab menu. To map one lab’s full ladder — its cheap, default, and top rungs and its reasoning mode — you’ll still open that lab’s own app or pricing page, since BoodleBox surfaces the specific models it offers rather than each lab’s entire pricing tier.

Reflection: Was the ladder obvious in the app, or hidden? Most consumer apps nudge you toward one default and bury the choice. Which rung is the app’s default, and is that the rung you would pick for most of your work?


Rep 2 — Match five products to their engines

List five AI features you actually encounter — the “summarize” button in your email, a customer-service chatbot on a website, the writing helper in your word processor, a search tool, a note-taker. For each, guess (or look up) which foundation model or lab is under the hood and whether you’re even allowed to choose. It’s fine to be uncertain — mark your confidence.

Reflection: How many of your five let you see or choose the underlying model? What does it mean for your budget and your privacy when a product hides which engine it’s using and where your data goes?


Rep 3 — Price a task three ways (paper exercise)

Using the Claude ladder in code/providers-and-families.txt (Haiku ~$1→$5, Sonnet ~$3→$15, Opus ~$5→$25 per 1M tokens, mid-2026), estimate the rough relative cost of running one task — say, summarizing a one-page note — 10,000 times on each tier. You don’t need exact math; you need the ratio. Then flag every number you wrote as a “mid-2026 snapshot.”

Reflection: Output tokens cost several times input. Which is bigger for a summarize task — the input (the document) or the output (the short summary)? How does that change which tier you’d pick for high-volume summarizing?


Reps 4–6: Feel the tradeoff

Rep 4 — Small vs flagship on a simple task

Pick a genuinely simple task: reword a two-line email politely, or turn five bullet points into a paragraph. Run it on a small/fast tier and on a flagship tier (or on two different assistants where one is clearly the “lighter” option). Judge honestly: can you tell which answer came from the bigger model without looking?

In BoodleBox. You don’t need two tabs for this: start one chat and @-mention two models — a lighter one and a flagship — so both answer the same prompt side by side. Type @ to open the bot picker and add each from the “AI Models” list. Fallback: run the same prompt in two public assistants and compare.

Reflection: Could you tell the difference? If not, what did paying for the flagship actually buy you on this task? End with your AI usage note.


Rep 5 — Small vs flagship on a hard task

Now pick a genuinely hard task with real steps: analyze the pros and cons of a decision with three competing factors, or work through a multi-part planning problem from your job. Run it on a small tier and a flagship (or a standard mode vs a reasoning mode). This time, look for where the weaker one breaks — a dropped constraint, a shallow answer, a logical slip.

In BoodleBox. Again, one multi-bot chat does the job: @-mention the small model and the flagship (or a reasoning-capable model) on the same hard prompt and read their answers next to each other — the side-by-side is what makes the flagship’s extra steps easy to spot.

Reflection: Where exactly did the smaller/standard model fall short — and would you have caught that slip if you weren’t comparing? Tie this to the spine rule: you own the verdict on both answers. End with your AI usage note.


Rep 6 — Toggle reasoning on and off

Take one shallow task (format a list) and one deep task (a multi-step logic or analysis problem). Run each with a reasoning/thinking/Pro mode on, then off. Note the difference in time and in quality for each.

In BoodleBox. Where a model exposes a thinking/Pro setting, toggle it; where it doesn’t, approximate the contrast by @-mentioning a reasoning-heavy model (a top Claude Opus or a Pro model, as of 2026) against a lighter one on the same task. Reasoning controls move fast, so describe the action you actually took in your usage note rather than a specific button.

Reflection: On which task did reasoning mode earn its extra time and cost — and on which was it pure waste? State the rule you’d give a coworker in one sentence. End with your AI usage note.


Reps 7–9: Turn the five dials

Rep 7 — Score three real tasks on the five dials

Pick three real tasks from your own week. For each, rate the five dials from Chapter 3.5: difficulty (1–5), speed need (low/med/high), budget sensitivity (low/med/high), privacy sensitivity (low/med/high), and multimodal? (does it involve images/audio/video?). Predict the recommended tier before you decide.

Task: _______________________________________________
  Difficulty (1-5): ___   Speed: ___   Budget: ___
  Privacy: ___   Multimodal? ___   ->  My tier guess: _______

Reflection: Which dial did the most work in your decision for each task? For most everyday professional work, is your honest answer higher or lower on the ladder than your instinct wanted?


Rep 8 — Run the Model Chooser and argue with it

Open the Model Chooser widget below Chapter 3. Enter each of your three tasks from Rep 7 and compare its recommendation to your prediction. Then change one dial per task and watch the recommendation move.

Reflection: Where did you and the widget disagree — and who was right, once you thought it through? (The widget is a helper, not an authority; you own the verdict.) Which single dial change flipped a recommendation most dramatically?


Rep 9 — Find the task where privacy overrides everything

Identify one real task involving information you should not paste into a public consumer chatbot (a client’s data, a patient detail, an unreleased plan — skim Appendix C). Set its privacy dial to high and note how that overrides the other dials — even if a cloud flagship is smarter, privacy sends you to a governed enterprise tool or a local model (previewed for Chapter 11).

In BoodleBox. For Concordia coursework and work, BoodleBox is that governed tool: as of 2026 it’s SOC-2 certified and doesn’t train on your data, which makes it an appropriate home for school/work content that a random public chatbot is not. (For truly regulated records — patient data, protected student records — still confirm the tool meets that specific obligation, and never paste more than the task needs.)

Reflection: Describe the task in a way that reveals nothing confidential. Why is choosing a “less capable” private option here the wiser choice, not the weaker one? Tie it to counting the cost (Chapter 3.8).


Rep 10: Seed the project

Rep 10 — Start your model-selection matrix

Open code/model-selection-matrix.csv in a spreadsheet app (Google Sheets, Excel, or Numbers — free options in Appendix A). Study the EXAMPLE row, then fill in two of the six task rows with real tasks from your job, turning all five dials and writing a one- or two-sentence justification for each on cost/speed/quality.

Reflection: Which of your two tasks was hardest to place, and why? Placing the ambiguous task well — the one that could go two ways — is the real skill this week is building.


Done? One Last Thing.

A miniature of Project 3, end to end. Take one real recurring task from your job. In a single page (a doc or a slide):

  1. Describe the task in one sentence — no confidential details.
  2. Turn all five dials and write the numbers/levels down.
  3. Name a provider and a tier you’d reach for — pin the model ID where you can (e.g., claude-sonnet-5), and date it “mid-2026.”
  4. Justify it in three lines — one each for cost, speed, and quality — plus a word on the privacy call.
  5. State the cost you counted: why not the flagship (or why it is worth it here).

If you can do this cleanly for one task tonight, you can do it for six in Project 3. That’s the whole job in miniature: turn the dials, choose the rung, count the cost, own the verdict.

Up next: Project 3 — Project P3: The Model-Selection Matrix.