Chapter 10 · Week 10

Creating with AI: Voice and Audio

The sheep know the shepherd's voice and flee a stranger's — whose voice is it, really, when it is cloned?

Chapter 10 — Creating with AI: Voice and Audio

“Trust, but verify.” — Russian proverb

“…the sheep follow him, for they know his voice. A stranger they will not follow, but they will flee from him, for they do not know the voice of strangers.” — John 10:4–5 (ESV)


Why This Matters

Last week you learned to make pictures and video that never happened. This week you learn to make something more intimate and, frankly, more dangerous: a voice. Not a robot reading a menu — a warm, breathing, emotionally-inflected voice that can sound exactly like a real human being. Including, if you feed it thirty seconds of audio, you. Including, if someone else feeds it thirty seconds of audio, your boss, your mother, your CEO.

Here is the change, stated plainly. For your whole life, a voice was proof. If the phone rang and it was your daughter’s voice saying she was in trouble, you did not ask her to prove it — the voice was the proof. Your ear was an authentication system, and it worked for a hundred thousand years. As of 2026, that system is broken. A convincing clone of a specific person’s voice can be built from a short sample by a consumer app in minutes, and played down a phone line in real time. “I recognized their voice” is no longer evidence of anything. That single sentence is the most important thing in this chapter, and we will come back to it three times before we are done.

But this is not a chapter of fear. The same technology that enables the scam call also narrates an audiobook in your own voice for a blind colleague, dubs your training video into thirty-two languages overnight, and turns a stack of dense reports into a twenty-minute podcast you can absorb on your commute. Text-to-speech, voice cloning, audio overviews of your documents, and full songs generated from a text prompt — these are real, useful, professional tools, available today, most of them from a browser. You are going to learn what each one is, what it is good for, and where the wheels come off.

And there is a second thread running under all of it, the one this whole book turns on: you choose the tool, and you own the verdict. The AI generates the voice; it cannot consent for the person the voice belongs to. The AI clones the sound; it cannot decide whether cloning it is honest. The AI writes the song; it cannot tell you whether you may legally sell it. Generation is cheap and instant now. Judgment — consent, disclosure, verification, accountability — is the scarce thing, and it is entirely yours. AI is an assistant here, not an authority, and in no domain is that line sharper than the human voice.

This week’s question is older than any microphone. In John’s Gospel, the sheep know the shepherd’s voice and flee a stranger’s — the voice is a mark of a person, a bearer of trust and presence. So when a voice is copied, sound-perfect, and sent out to speak words its owner never said: whose voice is it, really? Hold that. We will earn it by the end.

Coach’s Note — If this is your first chapter with me, the house rule is simple: learning is a sport, and the keyboard — here, the phone and the browser — is the gym. You do not learn to hear a fake by reading about fakes. Every section points at something to do. The reps and the project are where Week 10 actually gets into your hands. Show up.


10.1 — Why Your Ears Are No Longer Enough

Let us be precise about what changed, because the professional stakes hide in the details.

Synthetic voice is not new. Your GPS has talked to you for twenty years. What is new is the quality and the specificity. Two thresholds got crossed, roughly between 2023 and 2026:

  1. Indistinguishability. The best text-to-speech no longer sounds “computery.” It breathes, pauses, emphasizes, and carries emotion. In a blind test, most people cannot reliably pick the synthetic voice from the human one. The tell-tale robotic flatness is gone.
  2. Cloning from tiny samples. You no longer need a studio and hours of recording to copy a specific voice. Consumer tools can produce a usable instant clone of a particular person from well under a minute of audio, and a high-fidelity clone from a few minutes.

Put those together and you get the professional reality of 2026: any voice that exists as a recording anywhere — a podcast, a voicemail greeting, a conference talk on YouTube, a thirty-second Instagram clip — can be copied and made to say anything. That is a superpower for accessibility and a loaded weapon for fraud, and it is the same tool.

The lesson is not “audio is evil.” The lesson is that a category of trust you have relied on your whole life — I know that voice — quietly stopped being reliable, and most people have not noticed yet. You are going to notice, this week, and you are going to build the habits that a noticing professional keeps.


10.2 — Text-to-Speech: The New Baseline

Start with the workhorse: text-to-speech (TTS) — you give it written words, it gives you spoken audio. Conceptually it is the mirror image of the transcription you met in the productivity chapters. You type; it talks.

The reference tool professionals name most often as of mid-2026 is ElevenLabs — known for lifelike, emotionally aware voices, support for roughly 32 languages, and a large library (hundreds of prebuilt voices). Its Reader feature turns PDFs, ePubs, and articles into audio you can listen to, including some officially licensed well-known voices. It is not the only one: OpenAI’s ChatGPT Advanced Voice does real-time spoken conversation, Google and others ship strong voices, and specialist engines (Hume, Cartesia, PlayHT, and others) compete on naturalness and latency. Pricing is the usual moving target — most offer a free tier plus paid plans; check the vendor’s page rather than trusting a number I print here, because it will have changed by the time you read this.

Where TTS earns its keep for a working professional:

UseWhat it replacesThe win
Narrating a training video or courseBooking a studio, re-recording every editChange the script, regenerate in seconds
AccessibilityA colleague who cannot read the screenAny document becomes audio, instantly
Dubbing into other languagesA translation studioOne video, many languages, overnight
Voicemail / phone-tree greetingsRe-recording when details changeUpdate the text, not your voice
Turning a report into a listen-on-the-commute fileReading it at your deskTime you did not have, reclaimed

Notice that most of these are your own content, spoken in a generic or licensed voice. Nobody’s identity is at stake. That is the safe, boring, enormously useful center of this technology — and it is where you should spend most of your time. The excitement is in cloning. The value is mostly here.

Coach’s Note — The single most under-used professional application of TTS is accessibility, and it is the one I most want you to try. If a document you send can also be heard, you have widened who can receive your work — a colleague with low vision, a teammate with dyslexia, anyone whose commute is their only quiet hour. That is not a gimmick. That is your work reaching a neighbor it could not reach before.


10.3 — Voice Cloning, and the One Rule

Now the sharp tool. Voice cloning copies the timbre, cadence, and character of a specific person’s voice so that new text can be spoken as if by them. Two flavors:

  • Instant clone — from a short sample (seconds to a minute). Good, not perfect; fine for a personal narration or a rough draft.
  • Professional / high-fidelity clone — from several minutes of clean audio and (in reputable tools) a consent step. Close enough that listeners cannot tell.

The capability is genuinely useful and genuinely yours to use — on one condition, and I am going to put it in its own line because a validator, a lawyer, and your conscience will all check it:

You may clone only a voice you have consent to clone.

Your own voice: consent is trivial — it is yours. Anyone else’s voice — a coworker, a client, a celebrity, a politician, your late grandfather — requires their explicit, informed, documented permission, and there is no exception that starts with “but it would be so useful if…”. This is not merely etiquette. It is the seam where reputable tools, the law in a growing number of places, and plain honesty all agree. Reputable cloning tools now require you to attest that you have the right to clone a given voice, precisely because the temptation to skip that step is so strong.

This is the chapter’s spine rule made concrete. The tool will happily clone anyone — it has no conscience, no knowledge of your relationships, no standing to say no. The tool generates; you decide whether it is honest. Consent is a human judgment about a human relationship, and it can never be delegated to the software that stands to benefit from a “yes.”

We ship you the instrument that makes this real: a fill-in code/voice-consent-form.txt — a plain-language consent and authorization template covering whose voice, what uses are permitted, what is forbidden, how the clone will be disclosed, how long it is kept, and how consent can be revoked. You will sign it for your own voice in this week’s project. Read it now. Notice how much judgment lives in a form that a machine could never fill out for you.


10.4 — Audio Overviews: Your Documents, as a Podcast

Here is the one that surprises professionals the most. Google’s NotebookLM Audio Overviews takes a set of documents you upload — reports, a contract, research papers, meeting notes, a policy binder — and generates a two-host, podcast-style audio discussion of them. Two synthetic voices talk through your material as if they had read it and were explaining it to you on a walk.

It is uncanny and, for the right task, wonderful. Its distinguishing virtue is that it is source-grounded: it discusses your documents, not the open internet, and NotebookLM as a whole cites back to your sources. So the audio overview is a way to absorb a stack of dense material passively — on a commute, at the gym, while cooking — instead of reading forty pages at your desk.

Know its shape, as of mid-2026, before you lean on it:

  • Format is pre-styled. The two-host chatty podcast format and the voices are largely fixed — you do not get to fully configure them. It is a format, not a blank canvas.
  • Caps exist. The free tier allows only a few generations per day (roughly three); NotebookLM Plus, bundled into Google’s AI Pro plan (around $19.99/month as of mid-2026, ⚠ verify), lifts that substantially (roughly twenty). Numbers move — check Google’s page.
  • It can still be wrong. “Grounded in your sources” reduces hallucination; it does not eliminate the risk of oversimplifying, mis-emphasizing, or smoothing a nuance into a confident summary. It is a study aid, not an oracle.

That last point is the whole book in miniature. An audio overview is a draft understanding of your documents, spoken beautifully. It accelerates your reading; it does not replace your reading where the stakes are real. If you are going to decide something on the basis of those forty pages — sign the contract, quote the number in a board meeting — you verify the specific claim against the specific source. The audio drafts your understanding; you own the verdict on what it means.

Coach’s Note — The trap with an audio overview is that it is pleasant. Two friendly voices agreeing with each other feels like understanding. Pleasantness is not accuracy. When the overview makes a claim you are about to act on, pause the podcast, open the source document, and find the sentence. If you cannot find it, you do not yet know it — you have only heard it.


You can now generate a complete song — lyrics, vocals, and instrumentation — from a text prompt. Type “an upbeat acoustic folk song about a small-town bakery’s grand opening,” and a finished track comes back. The two names to know as of mid-2026 are Suno and Udio.

  • Suno is widely regarded as the best all-round AI song generator, especially for vocals. Its flagship was v5.5 (released around March 2026), with a free tier running on an earlier model. It has added “Voices” and custom-model features.
  • Udio is praised for instrumental clarity and stereo quality. Note a real practical catch as of 2026: ecosystem lock-in — you generally cannot freely export and use tracks off its platform. Read the terms before you build anything on it.

For a professional, the use cases are obvious: a jingle, background music for a video, a custom song for an event, a mood piece for a presentation. And here is where I have to be careful with you, because the honest answer is “it’s unsettled,” and I will not fake certainty I do not have.

The copyright status of AI-generated music is genuinely open as of mid-2026. The lawsuits are real and ongoing. Suno reached a reported $500 million settlement with Warner in November 2025, while suits from Universal and Sony were still ongoing as of mid-2026. The core disputes — whether training on copyrighted recordings was lawful, and who (if anyone) owns the AI’s output — are not resolved. I am your coach, not your lawyer, and I am not going to tell you a generated track is safe to sell, because as of this writing nobody can tell you that with authority.

So the professional posture is: use AI music freely to learn, prototype, and play; before you attach it to anything commercial or public-facing, read the tool’s current terms and, if real money or reputation is on the line, ask someone qualified. Category-level truth ages well (“full AI songs exist; the legal fight is unresolved”); a confident claim about what is “legal” will age like milk. This is the spine rule again: the tool generates the song in seconds; the decision about whether you may use it is a judgment call that lands on you.


10.6 — The Weaponized Voice: Vishing and the End of Voice Authentication

Now the section that could save someone you love real money, so read it twice.

Vishing — voice phishing — is a scam call, and voice cloning has made it dramatically more effective. The old scam call needed a smooth talker and a plausible story. The 2026 version can sound exactly like a person you trust. Two patterns dominate:

  • The family-emergency call. A cloned voice of a child or grandchild, panicked, says they have been in an accident or arrested and need money right now. The voice is right. The panic is convincing. The money is gone before anyone thinks to verify. These scams are widely reported; the specific aggregate dollar figures you will see quoted are mostly vendor estimates, not audited numbers — so I will not print one — but the pattern is very real.
  • The executive / vendor-payment call. A cloned executive voice, or a whole deepfaked video call, directs an employee to move funds or change payment details. The best-documented case is the January 2024 Hong Kong incident in which a finance worker at the engineering firm Arup paid out roughly US$25.6 million after a video call populated by deepfaked colleagues, including a fake CFO. That one is well-documented and safe to cite; it was video, but the same cloning that faked those faces fakes voices down a phone line every day.

Here is the defensive doctrine, and it is not technical — it is procedural, which means anyone can adopt it today:

The voice is no longer the authentication. The channel is. You do not verify by listening harder. You verify out of band — through a second, independent channel the attacker does not control.

Concretely:

  • Call back a known number. Not the number that called you — a number you already had. Hang up, call your daughter’s actual phone, your CFO’s actual line.
  • Use a code word. Agree, in advance, on a family or team word that a real emergency call will include and a scammer cannot know. This is the single best defense for the family-emergency scam, and it costs nothing.
  • Require dual approval for money movement. No single voice on a single call moves funds or changes payment details. Two people, two channels. Make it policy.

Notice the shape of the fix. It is the same move the good shepherd’s sheep make: they do not authenticate the stranger by the sound; they know their shepherd by a relationship and a presence the stranger cannot counterfeit. Out-of-band verification is that instinct, turned into a checklist.

Coach’s Note — Have the code-word conversation with your family this week. I mean it — not as a metaphor, as a chore on your list. It takes five minutes at dinner and it is the highest-return thing in this entire chapter. The best security control in the AI era is often a human agreement, made in advance, that no clone can forge.


Three habits separate a professional who uses synthetic audio responsibly from an amateur who gets themselves or their employer in trouble. Learn them as a set.

1. Consent — the rule from §10.3. You clone only voices you have documented permission to clone. For your own voice, sign the consent form so the scope is explicit: what it may be used for, what it may never be used for, and how you can revoke it. Consent is not a one-time shrug; it is a bounded, revocable grant.

2. Disclosure — say when it is synthetic. When you publish or send audio in a voice — especially a cloned one, and above all your own cloned voice speaking words you did not personally record — tell the listener it is AI-generated. A single honest line does it: “This narration was produced with an AI voice, created with my consent.” Disclosure is how you use a powerful tool without deceiving the person on the other end. Undisclosed synthetic voice, even for a good purpose, quietly spends trust you cannot easily earn back.

3. Provenance — attach and expect the “nutrition label.” The industry’s technical answer to “is this real?” is a dual-layer stack you met in Chapter 9, and it applies to audio too:

  • C2PA / Content Credentials — cryptographically signed metadata recording how a file was made. Rich, but strippable (a re-record or re-upload can remove it). ElevenLabs is among the 2026 adopters attaching these credentials to generated audio.
  • SynthID (Google) — an invisible watermark baked into the audio itself. It survives more handling but carries less information. Complementary to C2PA, not a replacement.

Use them where you can; do not trust them absolutely. The presence of a credential is evidence a file was AI-made; its absence proves nothing, because the label may have been stripped. Provenance is a helpful signal, not a verdict — the verdict, as always, is yours.

For the full house rules on consent, disclosure, and what may never be pasted into a public AI tool, see Appendix C. For the tool-by-tool audio and voice reference, see Appendix B.


10.8 — Choosing the Right Audio Tool

Do not reflexively reach for one app. Match the tool to the job, the way you learned to in Chapter 5.

You want to…Reach forWatch out for
Narrate content in a generic/licensed voiceTTS (ElevenLabs, ChatGPT voice, others)Nothing much — this is the safe center
Make content accessible as audioTTS “reader” featuresGet the pronunciation of names/terms right
Speak in your own voice at scaleVoice cloning — your voice, with a signed consent formDisclosure; safeguarding the clone from misuse
Speak in someone else’s voiceVoice cloning only with their documented consentIf you cannot get consent, you cannot do it. Full stop.
Absorb a stack of your documentsNotebookLM Audio OverviewPre-styled format; verify claims before acting; confidentiality of what you upload
Add music to a video / event / promoSuno, UdioUnsettled copyright — read terms before commercial use
Prove or check whether audio is AI-madeC2PA / SynthID detection toolsAbsence of a label proves nothing

One quiet warning that spans the whole table: confidentiality. The moment you upload a document to an audio-overview tool or paste text into a cloud voice tool, it has left your building. Client data, PII, anything under HIPAA or FERPA, anything your employer calls confidential — know your organization’s policy before it goes into any of these. The productivity win is real; so is the disclosure risk. (Chapter 14 makes this a discipline; Appendix C has the short version.)


Below this chapter on the website you will find an interactive panel: the Consent & Disclosure Checklist. Go use it now. It is not decoration; it is the rep that wires this week’s judgment into your hands.

The panel walks you through five realistic synthetic-audio scenarios — cloning your own voice for a course intro; using a coworker’s voice for a “fun” birthday message; cloning a well-known public figure’s voice for a promo “demo”; making a NotebookLM overview of a confidential client file; getting a panicked call in your daughter’s voice. For the first three, you run the four non-negotiable gates from §10.7 — written consent, disclosure to listeners, staying within the consented scope, and no deceptive impersonation — and the verdict banner turns green only when every one of them is honestly checked (two more boxes, keeping provenance attached and guarding the clone itself, are the recommended professional extras). The last two scenarios spring a different kind of trap: the checklist vanishes, because they were never consent questions at all — one is a confidentiality call you make before anything gets uploaded, and the other is an inbound scam you defeat by verifying out-of-band. For each scenario, the panel explains why — which rule applies, what the honest move is, and where the trap is hiding.

What it teaches is not a list of five specific answers. It teaches the reflex: the small internal pause, before you hit “generate” or “send” or “wire the money,” where you run consent → disclosure → verification. Until that pause is automatic, the rules in §10.7 are just words you read. After the checklist, they are a habit you own. Run it twice — once cold, then again after you have drafted your project’s ethics memo, and notice how much sharper your second pass is.


10.9 — Whose Voice Is It, Really?

Now the week’s question, given its due. When a voice is cloned — sound-perfect, indistinguishable, speaking words its owner never spoke — whose voice is it, really?

Jesus says of the good shepherd: “the sheep follow him, for they know his voice. A stranger they will not follow, but they will flee from him, for they do not know the voice of strangers” (John 10:4–5, ESV). Sit with what the sheep are doing. They are not authenticating the shepherd by acoustic analysis — pitch, timbre, waveform. They know his voice because they know him: the one who leads them out, who goes before them, who has kept them. The voice is not a sound to be matched; it is the mark of a person they are in relationship with. The stranger may learn the words; he cannot become the shepherd. And the sheep flee, because a voice detached from its true person is exactly what they will not follow.

That is the theological knife under this week’s technology. A cloned voice is sound without self — the acoustic signature severed from the person who is its rightful owner. It can carry a stranger’s intent while wearing a familiar sound, which is precisely the vishing scam and precisely what the sheep refuse. So “whose voice is it, really?” has a clean answer: the sound belongs to whoever generated it, but the voice — the mark of a person, the bearer of that person’s word and trust — cannot be transferred by a copy. You can steal the sound. You cannot steal the self. And to send out a stolen sound as though it were the self is to do exactly what the Eighth Commandment forbids: to bear false witness against your neighbor. Luther’s explanation is bracingly practical here — we are not merely to avoid lying, but to “defend him, speak well of him, and explain everything in the kindest way.” A voice clone deployed to deceive is the opposite motion: it borrows a neighbor’s most personal signature to speak a lie in their name.

This is why consent is not bureaucratic friction but the moral center of the whole chapter. To clone a voice with consent is to receive the sound as a gift, on terms the person set — bounded, disclosed, revocable. To clone it without consent is to take what was not given and speak in a name that is not yours. The form in code/ is small; the thing it protects is not.

And there is a hope buried in the passage for a noisy, counterfeit-saturated age. The sheep are not left defenseless against strangers’ voices — they have a Shepherd whose voice they truly know, and a knowing that no counterfeit can fool, because it rests on relationship and not on surface. The out-of-band verification you learned in §10.6 — the call-back, the code word, the second channel — is a faint secular echo of that deeper security: trust anchored not in the sound but in the relationship the sound is supposed to represent. The professional lesson and the spiritual one rhyme. In a world where any voice can be faked, you do not trust harder by listening harder. You verify through what the counterfeit cannot reach. A stranger they will not follow. Neither should you.


10.10 — Common Pitfalls

Pitfall: Cloning a voice you do not have consent to clone. Example: You clone a well-known local expert’s voice for a promo because “it’s just for a demo” and it would be so effective. Fix: The rule has no exceptions: clone only a voice you have documented consent to clone. Your own voice, yes — with a signed code/voice-consent-form.txt that names the scope. Anyone else’s, only with their explicit permission. If you cannot get consent, you cannot do it.


Pitfall: Treating a voice on the phone as proof of who is calling. Example: A panicked call in your grandchild’s voice asks you to wire bail money now, and you do — because the voice was unmistakable. Fix: The voice is no longer authentication. Verify out of band: hang up and call a number you already had, use a pre-agreed code word, and require dual approval for any money movement. Set the code word up before you need it.


Pitfall: Publishing or sending synthetic audio without disclosing that it is synthetic. Example: You send donors an “audio update from the director” in the director’s cloned voice, with no note that it was AI-generated, and it later comes out. Fix: Disclose. One honest line — “produced with an AI voice, created with my consent” — protects the trust that silence would quietly spend. Even a good purpose does not license undisclosed deception.


Pitfall: Assuming provenance metadata is permanent proof. Example: You rely on a Content Credentials label to tell real from fake, and treat a file without one as therefore authentic. Fix: C2PA metadata can be stripped by a screenshot, re-record, or re-upload; SynthID watermarks survive more but carry less. Presence of a label is a signal; absence proves nothing. Provenance informs your verdict; it is not the verdict.


Pitfall: Uploading confidential material into an audio tool without checking policy. Example: You drop a client’s confidential file into NotebookLM to make a handy audio overview for your commute. Fix: Uploading is exporting. PII, PHI (HIPAA), student records (FERPA), or anything your employer calls confidential must clear your organization’s policy before it goes into any cloud audio tool. See Appendix C.


Pitfall: Assuming AI-generated music is yours to sell. Example: You put a Suno track under a client’s paid ad campaign and treat it as royalty-free. Fix: As of mid-2026 the copyright status of AI music is genuinely unsettled — active litigation, no clean answer. Prototype freely; before anything commercial, read the tool’s current terms and, if real money or reputation is at stake, ask someone qualified. Do not treat “it generated” as “I own it.”


Pitfall: Trusting a pleasant audio overview as if it were the source. Example: You quote a figure in a board meeting because the NotebookLM hosts said it so confidently. Fix: Grounded ≠ infallible. An overview can oversimplify or mis-emphasize. Before you act on a specific claim, pause the audio, open the source, and find the sentence. If you cannot find it, you have only heard it, not verified it.


10.11 — Reps

The work is in the exercises. The keyboard — this week, your phone and your browser — is the gym. A preview of what is waiting:

  • Turn a document into audio with a TTS tool and judge whether it is good enough to send — and who it helps.
  • Fill in your own consent form (code/voice-consent-form.txt) and clone (or high-fidelity narrate) your own voice from the provided script.
  • Spot the fake — listen to synthetic and real audio and write down what tipped you off, or admit honestly that it didn’t.
  • Make a NotebookLM audio overview of a small document set, then fact-check one claim against the source.
  • Run the vishing drill — design a code-word and call-back protocol for your family or team, and write the one-line disclosure you would attach to a synthetic clip.

This week’s AI usage discipline: you may use every tool named here, but each rep that involves AI ends with an honest one-line AI usage note — what you asked, what it got wrong or oversimplified, and what you verified yourself. You choose the tool; you own the verdict.

A short Check Your Reps quiz is embedded on this page, right under the chapter. Take it before you move on — five questions, grounded in exactly what you just read.


10.12 — This Week’s Project

Your project is P10 — “Your Voice, With Consent,” specified in Project 10. You will clone (or high-fidelity narrate) your own voice — consent only — under a signed consent form, produce a short clip from the provided script with a proper disclosure, and write a one-page ethics memo on consent, disclosure, and misuse boundaries.

At a high level: Normal tier builds the consent form, the disclosed clip, and the ethics memo. Medium tier adds a NotebookLM audio overview of a document set, fact-checked against its sources. Hard tier asks for a decision memo — a defended recommendation on whether and how your fictional organization should adopt voice cloning at all, drawing the line between permitted and forbidden uses. That memo is the part no AI can write for you, because it is a judgment about consent, trust, and accountability — and that is exactly what gets graded hardest.


10.13 — Coach’s Final Word

Here is what I want you to carry out of Week 10. The technology to make a voice — any voice — is now sitting in your browser, free or nearly so, and it is astonishingly good. That is a gift and a hazard in the same hand. It narrates for the colleague who cannot read the screen, and it calls your mother in your own panicked voice to steal from her. The tool does not know the difference. You are the difference.

So the work of this chapter was never really about audio quality. It was about the three habits that make you the trustworthy one in the room: consent — you clone only what you have the right to clone; disclosure — you tell people when a voice is synthetic; and verification — you trust the channel, not the sound. Models and apps will turn over; half the version numbers in this chapter will be stale by the time you read it twice. Those three habits will not. They are the judgment the machine cannot supply, and they are yours to keep.

And under it all, the shepherd’s voice. A voice is the mark of a person, a bearer of trust — which is exactly why a stolen one does so much damage and a consented one is such a gift. Receive voices as gifts, on the terms their owners set. Send out only what is honestly yours to send. Verify through what no counterfeit can reach. A stranger they will not follow. Neither should you.

Now go do the reps. The Consent & Disclosure Checklist is waiting right below this page, the consent form and script are in code/, and Project 10 is where it all comes together.

See you on Monday.


Up next: Read the exercises and do all of Week 10’s reps, then build Project 10 — Project P10: Your Voice, With Consent. Keep Appendix B (the AI Tool Directory) open for the audio/voice tools, and re-read Appendix C on consent, disclosure, and confidentiality. Then on to Chapter 11 — Frontier vs Local: Cloud Power and Private AI.

Interactive Lab — Week 10
Consent & Disclosure Checklist

The tool generates the voice; you own the verdict. Pick a scenario, then run the four non-negotiable gates before you hit generate or send. The banner goes green only when every required box is honestly checked.

The situation
STOP
Not clear to proceed
Coach’s read

Try: Run the “Coworker’s voice” scenario and try to reach a green verdict. You can’t honestly check the consent box — and no amount of disclosure fixes a missing yes.

Check Your Reps

Check Your Reps — Voice & Audio

Question 1 of 5
Your phone rings and it is your daughter's panicked voice saying she has been in an accident and needs money wired right now. What does the chapter say is the right move?
Why: The voice is no longer authentication, so you verify out of band through a second channel the attacker does not control — a call-back to a known number plus a code word set up in advance.
Question 2 of 5
A coworker suggests cloning a famous local radio host's voice for a promo because it would be really effective. Under the chapter's one rule, may you?
Why: The single rule of voice cloning is that you may clone only a voice you have documented consent to clone — there is no 'but it would be so useful' exception, and disclosure never substitutes for a missing yes.
Question 3 of 5
You send donors an audio update in your director's cloned voice, created with her consent. What does the chapter say you still owe the listeners?
Why: Consent and disclosure are separate duties: even a consented, well-intentioned synthetic voice must be disclosed as AI-generated, or you quietly spend trust you cannot easily earn back.
Question 4 of 5
A voice clip arrives with no Content Credentials (C2PA) label on it. What can you correctly conclude?
Why: Provenance metadata like C2PA can be stripped by a re-record or re-upload, so a credential's presence is a helpful signal but its absence proves nothing — the verdict stays yours.
Question 5 of 5
You generated a great backing track with an AI music tool and want to put it under a client's paid ad campaign. What is the chapter's guidance as of mid-2026?
Why: As of mid-2026 the copyright status of AI-generated music is open with active litigation, so you prototype freely but check the tool's terms and get qualified advice before any commercial use.
YOU FINISHED. NICE WORK.