Skip to content

How to Make Vertical TV with AI [First Pilot]

Vertical TV means 9:16 microdramas that hook in 3 seconds. Here's the hands-on AI workflow that actually ships a pilot without credit burn or face drift.

6 min readBeginner

Why does my Vertical TV attempt look like a lazy crop?

I hit this exact wall last month. I’d watched a dozen short dramas on my phone, got hooked by the cliffhangers, then tried turning an old landscape clip into “Vertical TV.” Simple crop. Black bars everywhere, faces cut off, captions covering mouths. The swipe came in under two seconds.

Vertical TV isn’t a rotated rectangle. It’s 9:16 built for how people actually hold phones – episodes often 60-90 seconds or 1-2 minutes, cold open on conflict, one emotional turn, and a last frame that makes the next tap feel mandatory. Live-action full series still land around $100,000-$200,000 per Marketplace reporting (2026). AI only changes the math if you direct, not bingo-prompt.

Phone in one hand, thumb ready to flee: that posture is the whole format. If the face isn’t readable in the first glance, nothing else you generate matters.

Here’s how I shipped a usable 60-second pilot without burning a month of credits or watching the lead’s face mutate by scene four.

Quick context: the format that owns the scroll

The Vertical TV app on Google Play lists 5M+ downloads (as of the store page checked for this piece) and freemium 1-2 minute portrait episodes. Viewers watch muted a lot, so burned-in captions matter. Owl & Co put non-China vertical video near $150B in 2026 – treat that as a market snapshot that may have shifted – but the working rules stay narrow: one location or tight action per beat, faces big, no wasted establishing shots.

Aaj Tak’s Vertical TV does not simple-crop the news. Coverage of the launch describes AI real-time adaptive reframing plus pixel reconfiguration trained on 5+ years of viewer data so anchors, graphics, and tickers survive the jump from horizontal live feeds (BestMediaInfo). News and fiction share one job: the system has to pick what matters in the frame.

Hands-on: build your first Vertical TV pilot

I started free. Character App’s free Studio hands out Lite Images daily, so you can lock a Cast before any video spend. Rule I kept: no video credits until story and faces are nailed.

1. Write the 60-second spine first

Don’t open a video model. Open any LLM and force this structure:

Genre: modern revenge-romance, PG-13
Length: ~60 seconds screen time
Hook (first 3s): conflict or question on screen
Turn (middle): one reversal that changes everything
Button (last frame): maximum tension, no resolution
Scenes: 4-6 max, each one action / one location
Output as beat list only

My pilot: she opens a bank app and sees the transfer her partner made. Face falls. Door opens behind her. Cut on her turning. Done. No backstory dump.

2. Lock the Cast with multi-angle stills

Generate or upload a permitted photo. Front, three-quarter, side, neutral + two emotions, same wardrobe. Save them as sacred. One front-face reference is how drift starts – AXIS AI STUDIOS production notes (2026) keep flagging incomplete packs as the reason facial proportions and wardrobe wander after a few episodes.

Pro tip: light the reference pack the way your story will be lit. Warm interior refs fail hard under cool dramatic lighting later.

3. Direct scene-by-scene, draft cheap

In a tool that supports 9:16 natively and per-scene model choice, write plain-language action, assign Cast, set camera (“extreme close-up, face top third, quiet bottom for captions, slow push-in”). Draft the whole pilot on the cheapest/fastest model. Watch the assembly. Fix pacing. Only then upgrade the hook, turn, and cliffhanger.

Starter on Character App’s pricing page is $19.99/mo for 800 credits as of that page – roughly four minutes of default 5s clips if you believe their packing math. Draft-first is the difference between one solid pilot and an empty wallet. Credits show before you confirm; always read the quote. Text-only or other models differ.

  • Face / eyeline near the top-third line
  • Bottom third stays quiet for captions (most watch muted)
  • Singles over two-shots whenever possible
  • Hands out of extreme close-up or cut away – still a reliable tell

Export the linear 9:16 file. Light polish in CapCut or similar: hard captions, sting on the button, music under dialogue.

Common pitfalls that kill the pilot

Two burns, same week. I ran every scene at premium “just to see.” Credits gone before the cliffhanger. Then one pretty face ref plus mixed lighting notes – by scene three the lead looked like her own cousin.

Multi-character frames and reverse-angle backgrounds drift harder than singles. Cut between locked singles. Continuity checklist on every transition: face, hair, wardrobe color, prop in hand, screen direction. Adapting horizontal news or multi-window feeds with auto-reframe? Dense graphics and tickers still drop – public writeups describe the multi-layer need, not a failure-rate number. The model chooses what stays. It isn’t magic.

What the results actually look like

Finished pilot: 58 seconds clean vertical. Face stayed recognizable across four scenes because the pack was locked first. Paid spend stayed under one Starter month after free daily images for the Cast. Same length live-action would have eaten a crew day. Draft-upgrade is what kept spend under one month.

If the button doesn’t force the next tap, does pretty framing even matter? Retention is that blunt.

When NOT to use pure AI Vertical TV

Skip pure AI when you need perfect hand performance, complex multi-person blocking, or legal likeness of real public figures without clearance. Hybrid (AI backgrounds + real talent, or locked LoRAs where the stack allows) still wins for premium catalogs. Long-form prestige is the wrong bet – this format lives on snackable hits. Reframing live news with zero editorial pass? Expect mis-prioritized subjects on information-dense frames.

FAQ

Do I need a film background to make Vertical TV?

No. Three beats in plain English and a clear last frame. That’s enough to direct a pilot.

How many credits does a real pilot eat?

On a Starter-style plan, draft 4-6 short clips, then upgrade only the open, turn, and close – often a few hundred credits if you obey the in-app quote. Free daily images cover Cast iteration at zero video spend. Lock faces free → cheap assembly → upgrade three hero moments only.

Will the characters stay consistent for a full 70-episode run?

Only if the reference pack is production infrastructure, not a one-time prompt. Multi-angle, multi-expression, multi-lighting stills, used every generation session. After episode 3-5, incomplete packs are where facial and wardrobe drift compounds; two-shots and hand close-ups remain the loudest visual tells. No current consumer model makes long-run consistency automatic.

Open a free Studio, generate one locked Cast with today’s daily images, and write the four-beat pilot spine before you touch any video model. Ship the draft assembly tonight.