Here’s the mistake almost everyone makes when starting an AI video series: they treat episode 2 as a fresh generation. Same prompt, same character description, hit go. Then episode 3. Then episode 4. By episode 5 the main character has a different nose, a new jacket, and slightly wider eyes. The series looks like it was made by five different studios.
The models aren’t broken. You’re just using them like a photo generator. A series needs anchors – reusable visual reference points that pin identity across clips. Flip your workflow around that idea and the whole thing gets easier.
Key takeaway upfront
Build a character reference pack before you generate a single video clip. Feed those references into every clip generation as image inputs. Use the last frame of clip N as the starting frame for clip N+1. That’s the workflow – everything below is the detail.
Why series drift happens
AI video models generate each clip independently. They rebuild your character from whatever inputs you give – prompt, reference image, lighting, angle. Vidu’s consistency research confirms what most creators discover the hard way: one frontal face photo isn’t a reference set, it’s a single data point. The model guesses everything else from training data.
That guess drifts a little every clip. Small in one generation. Compounding across ten.
Prompt-only vs. anchor-frame: what actually breaks
Two ways to run a series.
| Approach | How it works | Where it breaks |
|---|---|---|
| Method A – Prompt-only | Write a detailed character description, reuse it in every clip’s text prompt | Face, wardrobe, and lighting drift by clip 3-4. No mechanism to correct it. |
| Method B – Anchor frames | Generate a reference sheet, feed images into every clip, use last frame of each clip as start of next | Coherence still fades after ~60s of chained runtime, but stays stable within a scene |
Method B wins. Not because it’s clever – because it removes the guessing.
The anchor-frame walkthrough (Veo 3.1)
Native audio, 4-8 second clips, 24 FPS – that’s Veo 3.1’s actual output window, per the Google Developer Blog announcement. It’s currently the strongest text-and-image-to-video option with synchronized audio, which is why this walkthrough uses it as the example.
Step 1 – Build the reference pack
Before touching a video tool, generate a still-image character sheet. You need at minimum:
- One clean frontal portrait (neutral expression, even lighting)
- One three-quarter angle
- One full-body shot showing the complete outfit
- Any signature props isolated on plain backgrounds
Save this pack. It’s the visual bible for every episode.
Step 2 – Generate clip 1 with image-to-video
Don’t start with text-to-video. Start from your reference image. In Google Flow or the Gemini API, select image-to-video, upload the frontal shot, and write a prompt that describes the action and camera move – not the character. The character is already defined by the image.
Prompt example:
"Medium shot, camera slowly pushes in. Subject looks left, then
turns to face camera and starts speaking. Warm afternoon light,
shallow depth of field. Ambient cafe sound, light chatter."
Notice what’s missing: no hair color, no outfit description, no age. All of that is in the image.
Step 3 – Chain, don’t restart
For clip 2, export the last frame of clip 1 and use it as the starting image for the next generation. This keeps continuity tight – the model connects a defined visual state to the next scene, rather than reinventing the character.
The catch: Inside Google Flow, the one-click “Extend” button has historically routed through Veo 2 Fast – you silently lose native audio and Veo 3.1 fidelity with no warning in the UI. Use the manual Frames-to-Video path instead: save the last frame, upload it as the start image for a fresh Veo 3.1 generation. Slower, but you keep the quality you paid for.
Step 4 – Do the pricing math before you commit
$2.00. That’s approximately what one 8-second 1080p Veo 3.1 Standard clip with audio costs via the Gemini API as of April 2026, according to BuildFastWithAI’s pricing breakdown. A 60-second episode is roughly eight chained 8-second clips – so ~$16 per episode in raw generation before re-rolls. Budget 2-3× that for the re-rolls you’ll definitely do. Veo 3.1 Lite costs about $0.05 per video (as of April 2026, per MindStudio’s comparison) – useful for draft storyboarding, but check the current Gemini API docs before committing, since pricing changes.
Edge cases nobody warns you about
Coherence decay at 60 seconds
Chain enough clips and the model starts forgetting original details. Community testing puts the practical breakpoint around 60 seconds of chained runtime – characters drift even with anchor frames after that point (source: glbgpt’s Veo 3.1 length guide). Fix: cut episodes into scenes and re-anchor from your original reference pack every ~45 seconds. The previous frame alone isn’t enough once you’re past that threshold.
Extension clips silently drop to 720p
Your first clip renders at 4K. The extension? Often 720p – Google’s extension pipeline downscales to save compute (per glbgpt’s testing). Eight crisp seconds, then noticeably softer for the rest of the episode. Workaround: upscale each extension separately, or accept 1080p as your ceiling and generate everything at 1080p from clip 1.
Veo 3.1 Lite is not a smaller Veo 3.1
Turns out the name is misleading. The official Google DeepMind model card confirms Veo 3.1 Lite is based on the older Veo 3 architecture – not 3.1. At ~$0.05/video it’s fine for previews, but audio quality and prompt adherence are older-generation. Use Lite for storyboarding, then re-generate the winning clips on Standard.
Sora is a dead end for series work
OpenAI announced the Sora web and app experiences were discontinued April 26, 2026, with the API following on September 24, 2026. Not a foundation worth building a channel on.
What a real episode actually costs you
Here’s the thing most pricing breakdowns skip: the math above assumes every generation works on the first try. It won’t. A 12-minute YouTube episode – the middle of the range where successful AI channels operate (per AI Magicx’s 2026 long-form guide) – means roughly 90 chained clips. At $2.00 each on Standard, that’s $180 before a single re-roll. With a modest 1.5× re-roll rate, you’re at $270 per episode. That’s not a complaint, it’s a planning number. Know it going in.
Is that high? Depends what you’d spend on a human production crew. But it does mean your reference pack isn’t just a quality tool – it’s a cost-control tool. Every clip you don’t have to re-roll because your character drifted is two dollars back in your pocket.
FAQ
Do I need paid Veo 3.1 to start an AI video series?
No. Prototype on Kling’s daily free credits or Google Flow’s trial tier. But native audio and the anchor-frame workflow described above only work properly on the paid Gemini API tier – so free gets you a proof of concept, not a finished episode.
How do I keep two characters consistent when they’re in a scene together?
Single-reference tools will fail you here. Vidu’s image-to-motion accepts up to seven reference images per generation – faces, costumes, props, backgrounds – keeping each entity visually consistent across the clip (as of 2026, per Vidu’s own documentation). Upload both character reference packs plus the shared environment as separate images. One photo per character still isn’t enough; you need the full three-shot pack (frontal, three-quarter, full body) for each person, otherwise the model averages them together in ways you won’t like.
Is text-to-video ever the right choice for a series?
Yes – for anything without a face that needs to match next episode. Landscapes, transitions, product close-ups. The moment a recurring character appears, switch to image-to-video.
Next action
Open your image generator of choice, pick one character concept, and produce the three-shot reference pack described in Step 1. Just the reference pack – no video yet. That’s the piece that determines whether your series holds together, and it’s the piece almost everyone skips.