The Model-Selection Matrix
Apologetic question: "What does it mean to count the cost — to value a tool rightly instead of always reaching for the most expensive one?"
Project P3 — The Model-Selection Matrix
“For which of you, desiring to build a tower, does not first sit down and count the cost, whether he has enough to complete it?” — Luke 14:28 (ESV)
Chapter: 3 — The Model Landscape: Providers, Families, and Tradeoffs Due: End of Week 3 Submit: A link to a shared folder (Google Drive / OneDrive / Dropbox) or a single PDF portfolio containing your completed matrix, your reflection, and — for the higher tiers — your screenshots and memo. Details under Submission below. No GitHub, no code. Allowed tools: BoodleBox (sign in with your Concordia account at box.boodle.ai — setup in Appendix A), where one chat’s “AI Models” picker lets you @-mention and compare the providers named below; fallback: the public assistants and their free tiers (ChatGPT, Claude, Gemini, Perplexity). Plus a spreadsheet app (Google Sheets, Excel, or Numbers); the tool directory in Appendix B. AI policy: You may use AI to research the landscape and to run the sample tasks. You may not ask an AI to fill in your matrix or write your justifications for you — the whole point is that you turn the dials and defend the call. Any place an AI helped, note it. You choose the tool; you own the verdict.
The Setup
You are the operations coordinator at Rivertown Family Dental, a nine-person practice: two dentists, three hygienists, a couple of assistants, a front-desk lead, and you — the person who keeps the whole thing running. You are not technical. You have never written a line of code and you never will. But last month the office manager handed you a small, real monthly AI budget and one sentence: “Figure out our AI tools and don’t blow the budget.”
Right now everyone at Rivertown just opens the same chatbot for everything, on personal free accounts, with no plan. Sometimes it’s great. Sometimes it’s slow and expensive for no reason. And once — this is the part that keeps you up — someone pasted a chunk of a patient’s record into a public chatbot to “help write a letter,” and you are fairly sure that should never happen again.
So you’ve decided to do what the tower-builder does: sit down and count the cost before you build. You are going to take the six tasks the office actually does with AI, and for each one decide — deliberately, in writing — which provider and which tier the office should reach for, and why, on cost, speed, quality, and privacy. When you’re done, you’ll have a one-page matrix any staff member can follow, and you’ll have earned the right to defend it to the office manager.
Here are the six real tasks to place (use these, or swap in six genuinely equivalent tasks from your own job):
- Appointment reminders & thank-you texts — short, friendly, routine, sent in high volume.
- Inbox triage — sort a day’s incoming general emails into “billing / scheduling / clinical question / spam.”
- The monthly patient newsletter — a warm, on-brand few paragraphs plus a headline.
- Reading a long vendor contract — a 40-page equipment lease; you need it summarized and to ask questions about specific clauses.
- Choosing between three vendor quotes — three offers with different prices, terms, and warranties; you need a reasoned recommendation.
- A letter that references a specific patient’s treatment history — genuinely confidential health information.
Notice that these are not all the same kind of task. That’s the point. Six tasks, and they do not all want the same rung of the ladder.
Setup (the starter)
This chapter ships two starter files in code/. You start from these — you do not submit them unchanged.
code/model-selection-matrix.csv— the blank matrix you will fill: one row per task, six scoring columns (the five dials plus volume), plus your recommended provider, tier, cost level, and justification. Open it in your spreadsheet app. One worked EXAMPLE row is included so you can see the shape.code/providers-and-families.txt— a one-page reference: the labs, the family ladder, the Claude price ladder (known-stable IDs), the five dials, and the open-vs-closed preview. Keep it open beside the matrix as you work.
In BoodleBox: as you decide each row, open the bot picker → “AI Models” to see which engines you can actually @-mention on campus (as of 2026: ChatGPT 5.6, Claude Opus 5, Gemini 3.1 Pro, Perplexity, image models). BoodleBox’s token-reduction keeps those runs affordable, which is part of why the honest answer for most rows sits low on the ladder. Fallback: the public assistants’ own model pickers.
Learning Targets
You will demonstrate that you can:
- Turn the five dials — difficulty, speed, budget, privacy, multimodal — on a real, mixed set of professional tasks.
- Match each task to a defensible tier, pinning a model ID where you can and dating every fast-moving fact as a “mid-2026 snapshot.”
- Count the cost — recognize when a cheap tier is not just acceptable but correct, and say so out loud instead of reflexively reaching for the flagship.
- Honor the privacy dial — recognize the one task where sensitivity overrides capability and route it away from a public consumer tool.
- Justify a choice in writing on cost, speed, and quality — the beginning of the professional judgment this whole course is built to grow.
Normal Tier
Goal: Fill the matrix for all six tasks with defensible, dated justifications.
Required features
- All six tasks scored. In
code/model-selection-matrix.csv, complete rows 1–6 (delete or keep the EXAMPLE row — your call). Every task gets a value in all six scoring columns (the five dials plus volume): difficulty (1–5), volume, speed priority, budget sensitivity, privacy sensitivity, and multimodal (Y/N). - A provider and a tier per task. For each, name a provider (or “any major assistant” where it genuinely doesn’t matter) and a tier (nano / mini / balanced / flagship / reasoning). Pin a model ID where you reasonably can (e.g.,
claude-sonnet-5,gpt-5.6) and mark whether a reasoning model is needed. - A justification per task. One or two sentences each, touching cost, speed, and quality — why this rung and not the one above or below it.
- At least one deliberate “count the cost” call. For at least one task, explicitly explain why you are not using the flagship — why a cheaper tier is the right choice, not a compromise.
- The privacy override, handled. Task 6 (the patient-history letter) must be routed away from a public consumer chatbot — to a governed enterprise tool (BoodleBox is Concordia’s vetted one: SOC-2 certified and it doesn’t train on your data, as of 2026) or a local model (previewed in Chapter 11) — with a one-line reason. For genuinely regulated records, confirm the tool meets that specific obligation and never paste more than the task needs. Skim Appendix C first.
- Everything dated and hedged. Any price, version, or “best” is marked as a mid-2026 snapshot. No confident permanent numbers.
- A short reflection (150–250 words,
.txt/.docx/.pdf): which placement surprised you, and where your first instinct was wrong.
Normal-tier rubric (out of 100)
| Criterion | Points |
|---|---|
| All six tasks scored in all six scoring columns (the five dials plus volume) | 20 |
| Each task assigned a provider + tier (model ID pinned where reasonable) | 15 |
| Justifications tie to cost, speed, and quality for every task | 20 |
| At least one deliberate “count the cost” call (why not the flagship) | 10 |
| Task 6 privacy override handled correctly (off the public chatbot) | 15 |
| Prices/versions hedged and dated; model IDs pinned, not marketing names | 10 |
| Reflection is honest and specific about a surprise or a corrected instinct | 10 |
| Total | 100 |
Medium Tier (+up to 25% extra credit)
Stop reasoning on paper and test one of your calls.
- Run one task on two tiers. Pick one task from your matrix. Actually run it on a small/fast tier and on a flagship tier (or a standard mode vs a reasoning mode). In BoodleBox, do this in one multi-bot chat:
@-mention both models on the same prompt so their answers sit side by side, then screenshot the comparison. Fallback: run it in two public assistants and screenshot each result. - Judge it honestly. In a short write-up (200–300 words), say whether the cheaper tier was “good enough,” where the pricier one was visibly better (if at all), and whether your Normal-tier choice for that task still stands or should change. Note the rough speed difference you felt.
- Add a cost column. Extend the matrix with a rough monthly cost estimate per task (order-of-magnitude is fine — “cents,” “a few dollars,” “tens of dollars”), and total it. Flag it as a mid-2026 estimate.
The grade here is in the honesty: if the flagship wasn’t worth it, say so; if it was, prove where.
Hard Tier (+up to 25% additional extra credit)
This is the judgment no model can produce for you, and it is graded as such.
Write ai-tool-policy-memo (one page, .docx or .pdf, addressed to “the office manager at Rivertown Family Dental”) that makes and defends a default model policy for the whole office. It must:
- State the default tier staff should reach for first, and the specific signals that tell them to climb (harder, longer, higher-stakes) or drop (simple, high-volume, cost-sensitive).
- Name the one non-negotiable rule about what must never be pasted into a public consumer tool, and the tool staff use instead for sensitive work.
- Name the cost you are accepting with your default — what you might give up by not defaulting to the flagship — and defend that as stewardship, not stinginess.
- Tie the recommendation explicitly to the spine rule — you choose the tool; you own the verdict — and to counting the cost (Luke 14:28).
A memo that says “always use the best model” or “always use the cheapest” fails. The grade is in the discrimination: showing you can tell a mini-tier task from a flagship task and write a rule a busy front-desk person could actually follow.
Submission
Put everything in one shared folder (Google Drive / OneDrive / Dropbox) or one PDF portfolio, and submit the link:
project-03/
model-selection-matrix.xlsx (or .csv / .pdf) <- your completed matrix
reflection.txt (or .docx / .pdf) <- Normal tier
tier-comparison.docx + screenshots/ <- Medium tier (if attempted)
ai-tool-policy-memo.docx (or .pdf) <- Hard tier (if attempted)
- Make the link viewable (set sharing to “anyone with the link can view”) so it opens without a request.
- Never put real confidential information in your submission. Describe Task 6 in a way that reveals nothing about a real person — use “a patient” and invented details only. This is itself part of the lesson (see Appendix C).
- Deliverables are documents and screenshots —
.xlsx,.csv,.docx,.pdf,.txt,.png. There is no code.
Hints
- Do the dials first, the tier second. If you pick a model before you’ve scored the task, you’re guessing. Score all five dials, then let the pattern point to a rung.
- Most of your six will land in the middle or bottom of the ladder. If every task somehow “needs the flagship,” you have not counted the cost — go back and be honest about which ones are actually simple.
- Only two of the six are genuinely hard. The long contract (Task 4, needs memory) and the vendor-quote decision (Task 5, needs reasoning) are the ones that earn a flagship or reasoning tier. The rest are mini/balanced work.
- Let privacy override. On Task 6, it does not matter that the cloud flagship is smarter. The right choice is the one that keeps the data home. Say that plainly.
- Pin IDs, date everything. Write
claude-sonnet-5, not “the newest Claude,” and tag every price “mid-2026.” When you’re unsure of a number, send yourself to the vendor’s pricing page rather than inventing one.
What Mastery Looks Like (Beyond the Rubric)
A mastered submission is one where I can read your matrix and your reflection and tell that you turned the dials — that you deliberately put the newsletter and the reminders on cheap rungs because the work didn’t need more, that you reached for a reasoning model on the vendor decision because it has steps, and that you never once let the patient letter near a public tool. The matrix is table stakes. The counting of the cost — choosing a humble tool on purpose and being able to defend it — is the project.
Coach’s Note — The tempting move this week is to hedge every task up the ladder so no one can say you “used a weak model.” Resist it. Anybody can spend the budget on the flagship; that’s not judgment, it’s fear wearing a lab coat. The professional who impresses me is the one who can look at a simple task and say, without flinching, “a mini model is the right tool here” — and mean it. That’s the whole skill.
When You’re Done (a short checklist)
- All six tasks have a value in all six scoring columns (the five dials plus volume).
- Every task has a provider, a tier, and a justification touching cost/speed/quality.
- At least one task explicitly explains why you didn’t use the flagship.
- Task 6 is routed off the public chatbot, with a reason.
- Every price and version is dated “mid-2026”; model IDs are pinned.
- Your reflection names a real surprise or a corrected instinct.
- (Medium) You actually ran one task on two tiers, with screenshots and an honest verdict.
- (Hard) Your memo defends a default policy a real coworker could follow.
- The shared link opens for anyone, and contains nothing truly confidential.
A theological footnote. The tower-builder in Luke 14:28 does one wise thing before he lays a single stone: he sits down and counts. He is not stingy — the tower still gets built — he is deliberate, refusing to begin what he has not weighed. This project is the smallest rehearsal of that wisdom in a modern office. The budget is your employer’s, not yours; the patient’s confidence is a trust, not a convenience; even the electricity a flagship burns to answer a trivial question is part of a creation you are stewarding. To reach reflexively for the biggest, most expensive tool is not diligence — it is the failure to count. To choose a humble model on purpose, because the task is small or the data is sacred, is not settling; it is faithfulness. Sit down first. Count the cost. Then build — and put your name on the verdict.
See you next week.