Chapter 09 · Week 9

Creating with AI: Image and Video

What does the commandment against false witness ask of us in an age of synthetic images?

Chapter 9 — Creating with AI: Image and Video

“Seeing is believing.” — English proverb

“You shall not bear false witness against your neighbor.” — Exodus 20:16 (ESV)


Why This Matters

Welcome to Phase 2. For eight weeks you learned to understand and choose AI — how the models work, how to pick one, how to prompt it, and, in Week 8, how to catch it when it lies. Now we turn the corner. Phase 2 is about making things. And the first thing we make is the thing that has changed the world’s relationship to truth faster than any other: synthetic images and video.

Here is what changed, stated plainly. For the whole history of photography, a photograph was a kind of witness. It was made by light bouncing off a real thing into a real lens. You could stage it, crop it, retouch it — but underneath there was, usually, a there there. The old proverb held: seeing was believing. As of 2026, that link is broken. A convincing, photoreal image of a scene that never happened, a person who never posed, a product that does not exist, can be generated in under a minute from a sentence you type on your phone. Short, realistic video — with synced audio — is now consumer-accessible too. The proverb at the top of this chapter is the thing this chapter takes apart.

That is a staggering creative gift and a genuine hazard, in the same tool, at the same time. The gift: a small-business owner with no budget can produce a professional-looking flyer; a teacher can illustrate a lesson; a nurse-educator can make a clear diagram; a marketer can iterate ten concepts before lunch. The hazard: the same button produces the fake CFO on the video call, the invented “customer photo,” the misleading before-and-after. This week you learn to use the gift and to carry the hazard responsibly. Those are not two skills. They are one.

Which brings us to the spine rule of this whole book, in its Week-9 form:

You choose the tool. You own the verdict. AI generates the pixels; you decide whether they are honest, whether you have the right to use them, and how you will label them. AI is an assistant, not an authority — and never less so than when it hands you something that looks exactly like the truth.

This week’s question is the oldest one about images, and Scripture stated it before cameras existed: what does the commandment against false witness ask of us in an age of synthetic images and video? “You shall not bear false witness against your neighbor” (Exodus 20:16, ESV). Hold that. A picture can bear false witness as loudly as a lying mouth — and now anyone can make one. We will earn that section by the end.

Coach’s Note — You do not need to be an artist to do this work, any more than you needed to be a novelist to write a good email in Week 4. Image generation is communication by description. Your professional skills — knowing what you want, knowing your audience, knowing what would mislead them — are exactly the skills that matter. The tool draws. You direct, and you decide.


9.1 — From Words to Pictures: What a Generative Image Tool Actually Does

Let us demystify the box before we drive it. A generative image tool takes your text description — the prompt — and produces a brand-new image that has never existed, pixel by pixel, to match it. It is not searching a library of stock photos and handing you the closest match. It is generating. Two people typing the same prompt get two different images; typing it again gets a third.

You do not need the math, but you need the right mental picture, because the wrong one leads to wrong expectations. The model learned, from an enormous number of image-and-caption pairs, the statistical relationships between words and visual patterns — what “morning light” tends to look like, how “terracotta” differs from “brick,” what “shallow depth of field” does to a background. When you prompt it, it starts from visual noise and, step by step, nudges that noise toward an image that fits your words. That is why it is fluent with look and feel and unreliable with facts: it knows what a keyboard looks like; it does not know that a keyboard has exactly the number of keys it has. Which is the same lesson as Week 8, in a new costume. It renders the plausible, not the verified.

Three consequences fall out of that, and they will save you a lot of confusion:

  1. It is non-deterministic. Re-running a prompt is not a bug fix; it is a fresh roll of the dice. To keep what you liked, you lock a seed or edit the image you have rather than re-rolling from scratch (more in §9.4).
  2. It is confidently wrong about details. Hands with six fingers, garbled text on a sign, a chair leg that melts into the floor, a reflection that does not match. As of mid-2026 the best tools have largely fixed hands and legible text — but largely is not always. You check.
  3. It has no idea what is true. Ask it for “a photo of our storefront” and it will invent a plausible storefront that is not yours. It cannot photograph a real thing. This is the single most important limitation for a professional, and the whole back half of this chapter is built on it.

Coach’s Note — If you came up thinking of “AI art” as the weird, six-fingered dreamscapes from a few years ago, update the picture. As of mid-2026 the mainstream tools produce images that routinely pass for real photographs to a casual viewer. That is precisely why the ethics half of this chapter is not optional garnish. The better the fake, the heavier the duty.


9.2 — The Image Toolbox: Who Wins Which Job

Just like general assistants (Week 5), there is no single “best” image tool in 2026 — there is the right tool for the job. Here is the honest lay of the land, as of mid-2026. Versions and prices move monthly; pin what you actually used and date it, and confirm on the vendor’s page. The fuller reference is Appendix B.

ToolMakerBest at (mid-2026)Watch for
GPT Image 2 (in ChatGPT)OpenAIThe default all-rounder: realism, following the prompt, legible text in images, editing. Replaced the older DALL·E line (~Apr 2026).Check plan/terms for commercial rights.
Midjourney (V8.1 default, ~June 2026)MidjourneyArtistic quality and stylization — the most beautiful, “designed” look; now does video too.Its own app/Discord; aesthetic can override realism.
Adobe Firefly (Firefly 4)AdobeCommercial safety — trained on licensed / public-domain content, offers IP indemnification, built into Photoshop; can run other models in one place.Sometimes a step behind on raw “wow.”
Flux (Flux 1.1 Pro)Black Forest LabsPhotorealism, fast, open-weight (can be self-hosted).You manage rights/hosting yourself.
Imagen (Imagen 4)GoogleReliable in-image text, strong quality; lives in Gemini.Tied to Google’s ecosystem.
Grok ImaginexAIImage + video in the Grok app; looser content rules.Looser rules cut both ways — more room to make something you shouldn’t.

Note the entry that is missing as a current product: DALL·E. For years it was OpenAI’s image model and the name everyone knew. As of mid-2026 it is legacy — superseded inside ChatGPT by GPT Image 2. If a colleague says “just use DALL·E,” they mean “use the image generator in ChatGPT,” which is now a different, better model. Names age; that is the point of dating everything.

Read that table for the pattern, not the leaderboard. Three durable lessons hide in it: (1) commercial safety is a real, buyable differentiator — Firefly’s licensed data and indemnification exist precisely because businesses got scared about copyright (§9.6); (2) the tools now edit as well as generate — you can fix one flaw without re-rolling the whole image; and (3) legible text inside images finally works reasonably well, which used to be the tell-tale sign of a fake. The leaderboard will churn. The pattern will not.

In BoodleBox (Concordia’s platform). You don’t need a separate account for each tool in that table. At CUW, your default place to generate images is BoodleBox — sign in with your Concordia account at box.boodle.ai. In a chat, type @ to open the bot picker and mention an image bot: as of 2026 that’s GPT Image 2 (@gpt-image-2) — the same all-rounder in the table above — or Nano Banana Pro (@nano-banana-pro). Same prompt, same five levers (§9.3); one login, no per-tool sign-ups, and it’s the FERPA/SOC-2 tool Concordia vetted. Video is still a separate step — BoodleBox is image- and text-first as of 2026, so the clips in §9.5 come from an outside tool. (Fallback: if you can’t reach BoodleBox, any tool in the table works on its own free tier.)


9.3 — Prompting an Image Is Not Prompting a Chatbot

In Week 4 you learned prompting as communication: objective, context, role, examples. Image prompting is a cousin, not a twin. A chatbot prompt is a paragraph of intent. An image prompt is closer to a shot list a photographer hands a crew — a stack of concrete, visual specifications. Vague in, vague out; here it is visibly vague.

There are five levers you turn on every serious image prompt. This is the spine of this week’s lab, so learn them by name.

LeverThe question it answersWeak → Strong example
SubjectWhat is in the frame?“a plant” → “a young tomato seedling in a terracotta pot, soil visible on two hands”
StyleWhat kind of image is this?(unsaid) → “natural editorial photograph, not stock, not cartoon, not 3D render”
LightingHow is it lit?(unsaid) → “soft warm morning light, gentle shadows, no harsh flash”
CompositionHow is it framed?(unsaid) → “close-up, subject on the left, uncluttered space on the right for text”
Aspect ratioWhat shape is the canvas?(default square) → “4:5 portrait for a flyer” (or 1:1, 16:9, 9:16)

Stack those and you go from “a plant” — which gets you a random plant in a random style — to a directed, on-brief image. Here is the same request, weak and strong:

WEAK:   a person planting a tomato

STRONG: A natural, editorial-style photograph of a pair of hands
        gently potting a young tomato seedling into a terracotta
        pot, dark soil visible on the fingers. Soft warm morning
        light from the left, shallow depth of field, blurred green
        nursery background. Close-up composition, subject on the
        left third, clean uncluttered space on the right for text.
        Warm, approachable, honest feeling. 1:1 square.

Read the strong version again. Every clause is one of the five levers. Nothing is decorative. That is what “prompting an image” means — not magic words, not “masterpiece, 8k, trending,” but a clear, specific description of the picture you can already see in your head.

Aspect ratio deserves a special word, because non-designers skip it and then wonder why nothing fits. It is the shape of the canvas, written as width:height. 1:1 is a square (social posts). 16:9 is wide, like a TV or an email banner. 9:16 is tall, for a phone screen or a story. 4:5 is a gentle portrait that performs well on social feeds and flyers. Pick the shape for where the image will live — designing a wide image and then cropping it to a square throws away half your composition. Most tools take aspect ratio as a setting or as --ar 16:9-style text; the lab lets you feel the difference.

Coach’s Note — The single biggest upgrade most people can make is to name the style and the lighting out loud. Leave them unsaid and the model picks for you — usually a slightly plastic, over-lit “AI look.” Say “natural editorial photograph, soft window light,” and the same subject suddenly looks like it belongs in a real magazine. You are the art director. Direct.


9.4 — Say What You Don’t Want: Negative Prompts, Iteration, and Editing

You have described what you want. Now, the three moves that separate a first-timer from someone who ships.

Negative prompts — saying what to keep out. Many tools let you specify what you do not want to see: --no text, watermark, extra fingers, neon colors. This is how you banish the recurring gremlins — the garbled sign, the sixth finger, the stock-photo gloss. If a tool has no explicit negative-prompt field, you say it in plain words: “no text baked into the image, no logos, natural colors.” Negative prompting is not an afterthought; it is half of directing. A good brief says both “warm and local” and “not glossy, not corporate.”

Iteration — the real skill. You will almost never nail it on prompt one, and you are not supposed to. The professional move is a deliberate loop:

  1. Write your best prompt (all five levers). Generate.
  2. Look with a cold eye. What is off? Too dark? Too cluttered? Wrong mood? Text where you wanted blank space?
  3. Change one thing in the prompt and regenerate. One lever at a time, so you learn what each change does.
  4. Repeat until it fits the brief — or until you decide this tool can’t get there and you switch.

Keep every version. “v1 → v2 → v3, and here is the one line I changed each time and why” is not busywork; it is the evidence that you directed the result, and it is a graded deliverable in this week’s project. Naming what you changed is how you build the instinct.

Editing instead of re-rolling. The 2026 tools don’t just generate — they edit. The key techniques, in plain language:

  • Inpainting — select a region and regenerate only that part (“fix this hand,” “remove that stray pot,” “clear this corner for text”) while the rest stays put.
  • Outpainting — extend the canvas beyond the original edges (turn a square into a wide banner without losing the subject).
  • Reference / seed — feed the tool an image (or reuse a seed number) so a new generation keeps a consistent look — vital when you need three images that feel like one set.

The difference matters: re-rolling throws away what you liked and gambles on a whole new image; editing keeps what works and fixes what doesn’t. When you have “almost the right image,” reach for editing, not the dice.


9.5 — Generative Video: The Capability Leap and Its Limits

Everything above scales to video now, and this is the frontier that moved fastest. As of mid-2026, consumer-accessible tools generate short, photoreal clips — often with synced audio — from a text prompt or a still image. That sentence would have been science fiction three years ago. The honest landscape (mid-2026; churny — hedge every version and price, confirm before you rely on it):

ToolMakerKnown for (mid-2026)
Sora 2 / Sora 2 ProOpenAIPhotoreal clips, strong physics, native audio. ⚠ Note: OpenAI has said the standalone Sora app/site is being wound down — confirm how Sora is offered before you rely on it as a product.
Veo 3.1GoogleStrongest all-rounder: follows the prompt, native audio, up to 4K; lives in Gemini/Flow. Rough usage ~$0.15/sec in fast mode ⚠.
Runway (Gen-4 / Gen-4.5)RunwayPro creative control — camera moves, motion brush, character consistency across shots.
Kling 3.0KuaishouMulti-shot cinematic sequences, subject consistency; ~$0.10/sec ⚠.
Seedance / Pika / othersByteDance / PikaUnified audio+video with lip-sync; consumer-friendly short clips.

Now the limits, because a professional plans around them, not around the demo reel:

  • Length. Think seconds, not minutes. These are clips — a few seconds of b-roll, a short product motion, a social snippet — not a two-minute explainer in one shot. You assemble longer pieces from many clips.
  • Physics and consistency wobble. Objects can morph, faces drift between shots, a hand passes through a cup, water behaves strangely. It has improved fast; it is not solved.
  • Fine control is limited. You cannot always get “her left hand rises exactly as she says the word ‘welcome.’” Tools like Runway add camera and motion controls, but precise choreography is still hard.
  • Cost and time. Video costs far more than images — per-second pricing, longer render waits, many re-rolls to get one usable clip. Budget accordingly.
  • The rights and disclosure stakes are higher. A fake photo can mislead; a fake video with audio can impersonate. Everything in §9.6–§9.8 applies to video doubly.

The practical professional posture in 2026: video generation is superb for short b-roll, concept mock-ups, illustrative motion, and social clips — and it is not yet your tool for anything that must be precise, long, or presented as a real recording of a real event. Know which job you are doing.

Coach’s Note — The demos you see online are the best-of-hundreds. The vendor rolled the dice fifty times and posted the winner. Your third roll will look rougher. That is not you doing it wrong; that is the honest hit rate. Plan for iteration and cost, and you will not be disappointed.


9.6 — Who Owns This? Rights, Likeness, and Commercial Use

Here the chapter turns from can I make it? to may I use it? — the question amateurs skip and professionals answer first. This is not legal advice, and the law is genuinely unsettled in 2026; it is the set of questions you must ask before an AI image goes out under your name or your employer’s. Deeper treatment lives in Appendix C.

Commercial use rights vary by tool and by plan. “I generated it, so I own it” is not automatically true. Whether you may use an AI image commercially — in an ad, on a product, in paid marketing — depends on that tool’s terms and often on whether you are on a free or paid tier. Some tools grant broad commercial rights; some restrict them; some are murky. Before you build a campaign on an image, read the terms for the specific tool and plan you used, and keep the receipt (tool, version, plan, date).

The training-data lawsuits are unresolved. As of mid-2026, multiple lawsuits are working through the courts over whether training image models on copyrighted work without permission was lawful. Nothing here is settled. Present it to yourself and your clients as an open question, not a decided one — and don’t give anyone legal comfort you are not qualified to give.

Indemnification is the market’s answer to that fear — and a real differentiator. Adobe Firefly trained on licensed and public-domain content and offers enterprise customers IP indemnification (Adobe will stand behind approved commercial use). For a business that cannot absorb a copyright risk, “a tool that promises to back us up” is worth more than a tool that looks slightly better. That is a business judgment — exactly the kind the machine can’t make for you.

Likeness is its own hard line. Generating a recognizable real person — a named celebrity, a public figure, or a private individual — raises rights of publicity and consent that have nothing to do with the image tool’s terms. You generally may not generate a real, identifiable person to sell something without their consent. “It’s AI, not a real photo of them” is not a defense; you made their likeness. When your image depicts a specific human being, the question is not “is this good,” it is “do I have consent.” (This is also next week’s whole subject, for voices — Chapter 10.)

Coach’s Note — The professional habit is boring and it saves careers: before you publish an AI image commercially, spend two minutes on three questions — May this tool’s terms let me? Does it depict a real person without consent? Is anything in the frame trademarked? If any answer is “not sure,” you stop and check. The code/image-rights-checklist.txt turns that habit into a worksheet.


9.7 — Provenance: The Nutrition Label on a Synthetic Image

If anyone can make a convincing fake, the world needs a way to ask “how was this made?” — a chain of custody for pixels. The industry’s 2026 answer is a dual-layer provenance stack, and you should understand both layers and their limits, because you will be asked about them.

Layer 1 — C2PA / Content Credentials: the visible “nutrition label.” C2PA (the standard) and its consumer face, Content Credentials, attach cryptographically signed metadata to a file recording how it was made and edited — which tool, when, what was AI-generated. Think of it as a nutrition label riveted to the image. It is rich and informative. Its weakness: metadata can be stripped. Take a screenshot, re-upload it, or run it through a tool that drops metadata, and the label can fall off. Present but fragile.

Layer 2 — SynthID and invisible watermarks: durable but thin. Google’s SynthID bakes an invisible watermark into the pixels themselves (also audio and text). It carries far less information than metadata — roughly “this was AI-generated” rather than the full history — but it survives many edits, crops, and screenshots that strip metadata. Durable but thin.

Why both. They are complementary. Metadata is rich but fragile; the watermark is durable but sparse. When OpenAI expanded provenance in May 2026, it framed the move explicitly as dual-layer — C2PA-conformant metadata and SynthID watermarks on images from ChatGPT, plus a previewed public verification tool. That is the shape of the whole industry’s answer: label loudly and watermark quietly, so that stripping one still leaves the other.

Adoption is real and spreading (mid-2026, ⚠ verify specifics before you cite them):

  • Google is adding SynthID detection in the Gemini app, with C2PA + SynthID detection coming to Google Search and Chrome; recent Pixel phones embed C2PA into video captures.
  • Meta / Instagram auto-apply Content Credentials labels to photos and videos.
  • Other 2026 adopters span Adobe (a long-standing C2PA backer), ElevenLabs, Nvidia, and more.

Here is the professional takeaway, and it is a judgment, not a checkbox. Provenance is a floor, not a ceiling. A file with a Content Credential telling you it’s AI is a gift — trust it. A file without one tells you almost nothing: absence of a label is not proof of authenticity, because the label may have been stripped, or the tool may never have added one. So provenance helps you believe the labeled far more than it helps you catch the unlabeled. Which is exactly why the human duty in the next section cannot be outsourced to a watermark.


9.8 — Deepfakes and the Duty to Label

Now the sharp edge. A deepfake is synthetic media — image, video, or voice — made to convincingly depict a real person doing or saying something they did not. The technology in this chapter, pointed at a real human, is the deepfake technology. Same button. Different target.

The canonical case, worth knowing by heart: in January 2024, in Hong Kong, a finance worker at the engineering firm Arup was tricked into paying out roughly US$25.6 million after joining a video call in which every other “colleague” was a deepfake — including a convincingly faked chief financial officer giving the instructions. The worker had doubts, saw familiar faces and heard familiar voices on the call, and was reassured. That is the whole lesson in one incident: “I saw them and heard them” is no longer authentication. Face-swapping and voice cloning have retired the oldest trust signal we have.

So what is the defense? Not sharper eyes — the fakes will outrun your eyes. The defense is procedure, out of band:

  • Verify through a second channel. Get a “CFO” video request to move money? Hang up and call back on a known number you already had — not one from the message. Confirm in person or on a channel the impersonator does not control.
  • Use pre-agreed code words for high-stakes requests among a team or family.
  • Require dual approval for money movement and sensitive actions, so no single deceived person can act alone.
  • Slow down. Urgency is the deepfaker’s favorite tool. “Right now, quietly, don’t tell anyone” is the tell, whatever the face on the screen.

(You will see per-incident and aggregate deepfake-loss figures in the trade press — ”>$200M,” and the like. Treat those as vendor estimates, not audited facts, and hedge them. The Arup case is the one that is well-documented enough to cite plainly.)

Now the other side of the same coin — not defending against fakes, but your duty when you are the one generating. This is where Exodus 20:16 stops being an epigraph and becomes a work instruction. The professional’s duties, in plain terms:

  1. Do not create synthetic media that impersonates a real person without their consent — full stop.
  2. Do not present a generated image or video as a real recording of a real event, product, or result. That is bearing false witness with a picture.
  3. Label synthetic media when a viewer could reasonably be misled. A caption — “AI-generated image” — costs you nothing and protects everyone. Keep provenance metadata intact where you can; add a plain-language disclosure regardless, because §9.7 taught you the metadata can fall off.
  4. Know your organization’s and platform’s disclosure rules — many now require labels — and follow them.

Labeling is not an admission of weakness. It is the mark of a professional who tells the truth about how their work was made. In a world where seeing is no longer believing, the person who says so is the trustworthy one.


9.x — Interactive Lab: Image-Prompt Builder

Below this chapter on the website you will find an interactive panel called the Image-Prompt Builder. Go use it now — it is where §9.3 gets into your hands.

The Builder gives you the five levers as controls: Subject, Style, Lighting, Composition, and Aspect ratio. You choose or type each one, and it assembles them into a single, well-formed image prompt — the kind you would paste into an image bot in BoodleBox (or into GPT Image 2, Midjourney, Firefly, or Flux directly). Start with a bare subject and watch the prompt stay generic; then add a style, a lighting choice, a composition, and a shape, and watch it turn into an art-director’s brief. Toggle the aspect ratio and notice how the same subject implies a different framing for a flyer, a square post, and a wide banner.

What it teaches is not a magic phrase — there isn’t one. It teaches the muscle of thinking in levers: that a good image prompt is a stack of deliberate visual decisions, each one of which you can name and defend. Build the three prompts for this week’s brief in the Builder first, then take them to a real tool for the project.

In BoodleBox: sign in at box.boodle.ai with your Concordia account, open a chat, type @ to open the bot picker, and mention an image bot — as of 2026, @gpt-image-2 or @nano-banana-pro — then paste each Builder prompt in. Do it twice: once fast and generic, once slow and specific, and compare what the two prompts produce. The gap between them is the whole skill. (Fallback: any free image tool works the same way — see Appendix A.)


9.9 — You Shall Not Bear False Witness

Now the week’s question, given its due. What does the commandment against false witness ask of us in an age of synthetic images and video?

“You shall not bear false witness against your neighbor” (Exodus 20:16, ESV). It is the Eighth Commandment, and for most of history it governed the mouth — the false accusation, the ruinous rumor, the lie told in court that costs a neighbor his freedom or his name. The commandment guards something precious: your neighbor’s standing in the eyes of others, which you have the power to wound with a word. Luther, explaining this commandment in the Small Catechism, presses it past mere honesty into active love: we should not “tell lies about our neighbor” or “hurt his reputation,” but should “defend him, speak well of him, and explain everything in the kindest way.” False witness is not only telling a lie; it is representing your neighbor to the world as something he is not.

Now put a camera — a generative camera — in that hand. For most of history the false witness had only words, and words are known to be fallible; a hearer can doubt them. An image was different. An image carried the authority of the eye: seeing is believing. That authority is exactly what synthetic media has seized. A generated picture or video borrows the trust we spent centuries placing in photographs, and spends it on something that never happened. This is why the Eighth Commandment lands with such force in 2026: it was written for exactly this power, long before the tool existed. To fake a person’s face, to stage a “photo” of an event that never occurred, to show a “result” a product cannot deliver — this is bearing false witness with a picture, and it wounds the neighbor precisely as the commandment feared: it deceives the one who sees, and it can ruin the one depicted.

But notice what the commandment does not say. It does not forbid making images. It does not forbid imagination, illustration, art, or a clearly-labeled picture of something that isn’t literally real. Scripture is full of made images that tell the truth — the parable that never happened yet reveals what is true, the tabernacle’s crafted cherubim. The sin is not fiction; the sin is false witness — passing off the unreal as real to deceive a neighbor to his cost. That is the line the professional walks. An AI image on a flyer, understood by all to be an illustration, bears no false witness. The same image presented as a real photograph of your real store, or of a real customer who never existed, does. The pixels are identical. The representation is what the commandment judges.

So the Eighth Commandment gives us the working rule for this entire chapter, and it is stricter than any platform policy: tell the truth about what you have made. Do not impersonate. Do not pass off the generated as the captured. And where a viewer could be misled, say so — label it, disclose it, put the best and most honest construction in front of your neighbor’s eyes. Luther told us to defend our neighbor’s reputation and speak well of him. In 2026 that includes not fabricating him, and not deceiving him. The disclosure line you write at the bottom of an AI image is not red tape. It is the Eighth Commandment, keyed into a caption box.

And this is where the spine rule and the Scripture say the same thing. You own the verdict because you are the witness. The tool cannot be a false witness — it has no neighbor, no oath, no account to give. You do. When the image goes out under your name, you have testified. Make it a true testimony.


9.10 — Common Pitfalls

Pitfall: Assuming “I generated it, so I can use it commercially.” Example: A café owner puts an AI image on paid ads, on a free tier whose terms don’t grant commercial rights — or from a tool now caught up in a copyright suit. Fix: Check the specific tool and plan’s commercial-use terms before you build on an image; keep the receipt (tool, version, plan, date). For work that can’t absorb risk, prefer a commercially-safe tool with indemnification (e.g. Firefly). Use code/image-rights-checklist.txt.


Pitfall: Presenting a generated image as a real photograph. Example: An AI “customer photo” or a “before/after” of a result the product can’t actually deliver, posted as if it were real. Fix: Never pass off the generated as the captured. If a viewer could reasonably think it’s a real photo of a real thing, either don’t use it or label it clearly. That’s the Eighth Commandment, not just marketing policy.


Pitfall: Generating a real, recognizable person without consent. Example: Making a fake image or clip of a named public figure — or a private individual — to promote something. Fix: Depicting a specific real person is a consent-and-likeness question that the tool’s terms don’t cover. Get written consent, or don’t depict them. “It’s AI, not really them” is not a defense — you made their likeness.


Pitfall: Trusting your eyes to spot a deepfake. Example: Approving a money transfer because the “CFO” on the video call looked and sounded right (the Arup pattern). Fix: “I saw and heard them” is no longer authentication. Verify high-stakes requests out of band — call back a known number, use a code word, require dual approval. Slow down when someone insists on urgency and secrecy.


Pitfall: Treating provenance (Content Credentials / watermarks) as a complete solution. Example: “It has no AI label, so it must be a real photo” — when the label was simply stripped by a screenshot. Fix: A label present is trustworthy; a label absent proves nothing. Provenance helps you believe the labeled, not catch the unlabeled. Keep your own provenance intact, and add a plain-language disclosure regardless.


Pitfall: Re-rolling the dice when you should be editing. Example: You get an almost-perfect image with one bad hand, so you regenerate twenty times and lose the version you loved. Fix: When you have “almost right,” use editing — inpainting, outpainting, reference/seed — to fix the one flaw and keep the rest. Re-roll only when the whole image is wrong.


Pitfall: Skipping style and lighting in the prompt. Example: “a plant in a pot” comes back plastic, over-lit, and generic — the classic “AI look.” Fix: Name the style (“natural editorial photo, not stock”) and the lighting (“soft warm morning light”) explicitly. Turn all five levers on every serious prompt (§9.3). You’re the art director; direct.


9.11 — Reps

The work is in the exercises. The keyboard — and this week, the image tool — is the gym. A preview of what is waiting:

  • Turn one vague prompt into a five-lever prompt and generate both, so you see the difference specificity makes.
  • Iterate on brief: take one deliverable from the code/image-brief.txt and drive it v1 → v2 → v3, changing one lever at a time.
  • Edit instead of re-roll: fix a single flaw with inpainting and keep the rest of the image.
  • Run the rights checklist code/image-rights-checklist.txt against a real generated image and write the verdict line yourself.
  • Find the Content Credential: generate an image, then check whether it carries provenance and whether a screenshot strips it.

This week’s AI policy for reps (Phase 2): you are creating with AI now — that is the point. But you still own the verdict. Every rep that generates media ends with an honest AI usage note: what tool and version you used, what you had to iterate, and — for anything you’d publish — the rights and disclosure decision. AI generates; you decide, verify, and label.

A short Check Your Reps quiz is embedded on this page, right under the chapter. Take it before you move on — five questions, grounded in exactly what you just read.


9.12 — This Week’s Project

Your project is P9 — “Generate on Brief,” specified in Project 9. You will take a real creative brief — code/image-brief.txt — and produce a small, on-brand set of images with an AI image tool: iterating the prompt through named versions, applying the five levers deliberately, and then doing the part that separates a professional from a button-pusher — documenting the rights and disclosure decisions with code/image-rights-checklist.txt.

At a high level: Normal tier produces the on-brief image set with documented iteration and a completed rights checklist. Medium tier adds a second variation set, an edit-not-re-roll fix, and a cross-tool comparison. Hard tier asks for a one-page recommendation memo to the business owner — which images they may publish commercially and under what disclosure, and which they must not use, and why. That memo is the judgment no image tool can generate for you, and it is where this week’s thesis gets graded.


9.13 — Coach’s Final Word

Here is what I want you to carry out of Week 9. You just picked up a genuinely astonishing tool — the ability to make, from a sentence, an image or a clip that looks like the truth. That power used to belong to studios and budgets. Now it is on your phone. Use it. Make the flyer, illustrate the lesson, mock up the concept, save the small business the money it doesn’t have. This is real professional leverage, and I want you fluent with it.

But you also picked up a duty, and it is exactly as large as the power. The old proverb is broken: seeing is no longer believing. Which means the trust that photographs carried for a century now rests on people — on whether the person who made the image is willing to tell the truth about it. That person is you. You choose the tool. You own the verdict on whether it is honest, whether you have the right to use it, and how you will label it. No watermark and no policy can carry that for you.

The models in this chapter will have higher version numbers by the time you read it twice. The five levers, the rights questions, the out-of-band verification, and the plain caption that says “AI-generated” — those outlast every product in the table. And underneath them runs the oldest instruction of all, written for exactly this moment: you shall not bear false witness against your neighbor. Make true images. Label the rest. Put your name on a testimony you’d defend.

Now go do the reps. The Image-Prompt Builder is waiting right below this page, the brief and the checklist are in code/, and Project 9 is where it comes together.

See you on Monday.


Up next: Read the exercises and do all of Week 9’s reps, then build Project 9 — Project P9: Generate on Brief. For the image and video tool directory see Appendix B; for copyright, likeness, consent, and disclosure see Appendix C; for account setup see Appendix A; and for terms like provenance, C2PA, and multimodal see the glossary in Appendix D. Then Chapter 10 — Creating with AI: Voice and Audio.

Interactive Lab — Week 9
Image-Prompt Builder

A strong image prompt is a stack of deliberate visual decisions — a shot list, not a sentence. Turn the five levers below (plus color/mood and a negative prompt) and watch a bare subject become an art-director's brief. Start generic, then add each lever and feel the prompt sharpen.

Tip Name the subject concretely. "a plant" gets you a random plant; "a young tomato seedling in a terracotta pot" gets you the picture in your head.
Your assembled prompt
Levers on:
Try: clear the Style and Lighting fields and read the prompt — that's the generic "AI look." Now put them back and toggle the aspect ratio between 4:5 and 16:9. Same subject, different brief. The gap between the two prompts is the whole skill.
Check Your Reps

Check Your Reps — Creating with AI: Image and Video

Question 1 of 5
When you give a generative image tool a prompt, what does it actually do?
Why: The tool generates a fresh image from visual noise guided by your words — it isn't retrieving stock photos, can't photograph real things, and is non-deterministic.
Question 2 of 5
You prompt "a plant in a pot" and get a plastic, over-lit, generic result. What's the fix the chapter recommends?
Why: The single biggest upgrade is naming the style and lighting out loud; leave them unsaid and the model defaults to the plastic "AI look."
Question 3 of 5
A café owner says, "I generated the image myself, so I can definitely use it in a paid ad." Is that right?
Why: "I made it, so I own it" isn't automatically true; commercial rights vary by tool and plan, so check the terms for what you actually used and keep the receipt.
Question 4 of 5
After the Arup case — a worker paid out ~US$25.6M on a video call where the "colleagues," including the CFO, were deepfakes — what is the real defense?
Why: "I saw and heard them" is no longer authentication; the defense is out-of-band procedure and slowing down, not sharper eyes.
Question 5 of 5
You receive an image that carries no AI or Content-Credential label. What can you conclude?
Why: Provenance is a floor, not a ceiling: a label present is trustworthy, but a label absent proves nothing because it may have been stripped or never applied.
YOU FINISHED. NICE WORK.