AI video models
Making AI storyboards reliable: five rounds of experiments with Seedream and Seedance
How we tested one-canvas storyboard sheets for AI filmmaking — prompt registers, sequential mode, vertical formats, and panel detection — with every result measured.
- AI storyboarding
- Seedream 4.5
- Seedance 2.0
- image generation reliability
- AI filmmaking

A storyboard is the connective tissue of a film: before anything expensive renders, you want six angles on a scene that agree with each other — same faces, same light, same geography. We're building a flow where a single generated storyboard sheet (a 2×3 grid of shots in one image) is cropped into individual panels, and the user clicks the panels they want as exact visual references for video generation with Seedance 2.0.
That only works if sheet generation is reliable: correct grid geometry, no sub-images inside a panel, six genuinely different setups, and one consistent cast across all of them. So before wiring it into the product, we spent five rounds of controlled experiments — roughly a hundred generations and under $10 of image spend — finding out what actually holds up. This post reports everything we measured. Sheets were generated with ByteDance's Seedream 4.5, with Google's Nano Banana Pro as a challenger in one round.
The results, up front
| Question | Verdict | Key evidence |
|---|---|---|
| One-canvas grid vs. separate images | Grid wins. Panels generated in one canvas share one attention context, which is what preserves identity | Character consistency 8.3–9.7/10 across every grid round |
| Best prompt register | "Photographer's contact sheet" + strict-uniform-lattice language | 97% pass on landscape (29/30) after iteration |
| Seedream's native sequential mode | Rejected — perfect geometry, but shots collapse into near-duplicates | 1/5 distinct; 6.4× slower; ~5× cost |
| Repeating character traits in every panel ("trait-locking") | Hurts. Describe the cast once, globally | Crop usability fell 4/5 → 2/5 with no consistency gain |
| Asking for drawn divider lines | Hurts | Pass rate fell across every gate |
| 2×2 instead of 2×3 | No reliability edge worth two fewer shots on landscape | 2×2: 5/5 crops usable vs. 2×3: 4–5/5 |
| Vertical (9:16) sheets | Cross-model failure ~50% — portrait canvases trigger comic-page layouts in both Seedream and Nano Banana Pro | Seedream 2/6, NB Pro 1/6 with identical prompts |
| Panel-boundary detection | The vertical fix. Crop the panels the model actually drew, accept whatever count came back | 73% of "failed" vertical sheets still yielded 4+ usable panels |
| User-directed per-panel descriptions | Works. Written shot direction lands in its assigned panel | 6/6 slot adherence, every repeat |
| Panel aspect ratio | Must equal the film's output ratio, with grid orientation following cell orientation | Mis-oriented grids measurably degrade layout discipline |
The rest of this post is how we got those numbers.
What good looks like: one generation, six camera setups, one consistent performer. Crowd scenario, contact-sheet register, Seedream 4.5, ~25 seconds.
How we tested
Every round used the same harness: generate sheets across deliberately hard scenarios (night-rain lighting, crowds, macro product shots, creatures, vertical formats, scenes with character reference images attached), crop them exactly the way production would, and grade with a vision-language-model judge against explicit gates — every crop is one complete coherent shot, no annotation text, six distinct camera setups — plus 0–10 scores for character, location/lighting, and style consistency.
Two methodology lessons earned the hard way, which we'd pass on to anyone running similar evaluations:
- Never trust the judge without looking at pixels. Our first reliability run scored 55%; visual verification showed the true rate was ~70%. The judge was failing sheets for diegetic text — a clipboard reading "FAILED INSPECTION" in a scene about a failed inspection, a painted "B3" in a parking garage set on level B3 — and for cosmetic gutters that don't affect crops at all. Rubrics need an explicit in-world-text exemption, stated twice.
- Repeats are the point. Grid failures are probabilistic. A scenario that passes once can fail the next run; single-sample comparisons measure luck.
Round 1: six approaches, head to head
We screened four prompt variants, a 2×2 board, and Seedream's native
sequential mode (sequential_image_generation, which returns a set of
separate images instead of a grid) across five hard scenarios.
The content quality surprised us everywhere — with character reference images attached, both of our test characters stayed on-model in every panel of every sheet. The differences were all in layout discipline and shot diversity:
- The contact-sheet register ("a photographer's contact sheet of 6 cinematic film stills…") beat the word "storyboard," which pulls in borders, panel numbers, and caption bars from the model's training prior.
- Trait-locking — repeating each character's full description inside every panel line, a popular community technique — actively hurt crop usability and bought nothing on consistency. The academic literature (One-Prompt-One-Story) predicted this: identity lives in the shared prompt context, and re-description invites drift.
- Sequential mode produced flawless full-frame geometry and collapsed creatively: five of six runs returned near-duplicate framings — six similar medium shots of the same subject rather than wide/medium/close coverage. At 160s versus 25s per board and ~5× the cost, it lost on every axis we care about except geometry, which round 5 solves more cheaply.
Sequential mode's failure, laid flat: same scenario as the sheet above, but the six separately-generated images collapse toward one framing. Consistency and shot diversity pull in opposite directions, and the set generator gives up the diversity.
Reference images carry hardest: both characters here are generated from attached identity plates, and hold across all six panels — down to the diegetic prop text on the inspection report, which our first judge rubric wrongly penalized.
The register generalizes beyond characters — product scenarios keep the same object, lighting, and grade across the coverage ladder. Macro-only scenes remained the hardest distinctness case through every round.
Rounds 2–3: from 55% to 97% on landscape
We locked the winning register and measured it properly: twelve scenarios, three repeats each. The corrected first pass sat around 70%, with three real failure modes: comic-page layouts (the model drawing merged and irregular panels — beautiful manga pages, useless for fixed-grid cropping), near-duplicate panels, and the occasional sub-composite (two little images inside one cell).
The fix was naming the failure modes in the prompt: every still identical in size, a perfectly regular lattice, NOT a comic page, NOT a manga page, no merged or spanning panels — plus an explicit distinctness ladder (at least one wide, one medium, one close-up). One iteration took the weak scenarios from 55% to ~90% overall, and landscape — the dominant film format — to 29/30 (97%), our ship bar. Square format held at 2/3 with the failure being a single weak panel, not a layout collapse.
Vertical did not follow, which became its own investigation.
Round 4: vertical is a model problem, not a prompt problem
On portrait canvases, both models kept fitting panels to the canvas rather than to the request — Seedream drawing full-width webtoon-style rows, Nano Banana Pro drawing immaculate grids with the wrong count (a clean 4×2 when asked for 3×2). With the identical anti-comic prompt, Seedream passed 2/6 vertical sheets and NB Pro 1/6. Conclusion: the tall-canvas layout prior is trained-in across vendors. You cannot prompt your way out, and switching between these two models doesn't help either.
Vertical failure shape one — Seedream's comic-page prior: a gorgeous manga layout that fixed-grid cropping slices to pieces.
Vertical failure shape two — Nano Banana Pro: flawless discipline, wrong count. A clean 4×2 when 3×2 was requested. Both sheets are full of good panels; only the grid assumption failed, which is what pointed us at round 5.
Round 5: stop asking, start detecting
The reframe that solved it: in almost every "failed" vertical sheet, the panels themselves were good — only our assumption about where they sat was wrong. So instead of demanding a rigid grid, we detect the gutter lines the model actually drew and crop the cells that exist, accepting whatever regular count came back. A sheet with eight good panels instead of six isn't a failure; it's two free extra angles.
Re-processing every vertical sheet from earlier rounds with a prototype detector: the wrong-count NB Pro sheets were fully recovered (8/8 usable panels each), and 73% of all vertical sheets yielded four or more usable, distinct panels — from a format that scored ~30% under blind cropping. The production version adds plausibility guardrails and falls back to the uniform grid for flush sheets, and doubles as a safety net for the rare landscape sheet with drawn gutters.
What shipped
The measured recipe is now our production configuration:
- contact-sheet register with uniform-lattice and anti-comic-page language, numbered per-panel lines (which users can override with their own shot descriptions — the 6/6 adherence result), one global cast description;
- sheet dimensions computed so each cell is exactly the film's aspect ratio, with grid orientation matched to cell orientation;
- crops with a small inner margin (every working grid pipeline we found — ours included — needs one to absorb boundary bleed);
- gutter detection with variable panel counts, so the flow keeps every good panel the model draws;
- and the panels feed Seedance 2.0 as ordered reference images, which honors their composition shot by shot.
Every number
For readers who want the full picture, the round-by-round results. "Crops usable" = every cropped panel is one complete coherent shot; "distinct" = six clearly different camera setups; "char" = judged character consistency (0–10). All generations Seedream 4.5 unless noted.
Round 1 — approach screen (5 hard scenarios × 1 repeat per arm):
| Arm | Crops usable | Distinct | Char | Mean latency |
|---|---|---|---|---|
| 2×3 · contact-sheet register | 4/5 | 5/5 | 9.0 | ~25s |
| 2×3 · baseline "storyboard" prompt | 5/5 | 3/5 | 7.8 | ~25s |
| 2×3 · trait-locking | 2/5 | 4/5 | 9.0 | ~24s |
| 2×3 · requested divider lines | 3/5 | 3/5 | 8.8 | ~23s |
| 2×2 · contact-sheet register | 5/5 | 4/5 | 9.0 | ~24s |
| Sequential mode (6 images) | 5/5 | 1/5 | 8.2 | ~160s |
Rounds 2–3 — reliability runs (12 scenarios × 3 repeats; pass = all product gates):
| Scenario | Format | First prompt | After lattice iteration |
|---|---|---|---|
| Two-hander w/ character refs (diner) | 16:9 | 1/1¹ | 3/3 |
| Concert crowd | 16:9 | 2/2¹ | 3/3 |
| Creature exterior | 16:9 | 3/3 | 3/3 |
| Macro product | 16:9 | 0/3 | 2/3 |
| Neon chase | 9:16 | 0/3 | 3/6² |
| Café two-hander | 1:1 | 1/3 | 2/3 |
| Directed shot list | 16:9 | 1/3³ | 3/3 |
| Period ballroom | 16:9 | 3/3 | 3/3 |
| Nature documentary | 16:9 | 2/3 | 3/3 |
| Two-character comedy w/ refs | 16:9 | 1/3³ | 3/3 |
| Kitchen drama | 16:9 | 1/3 | 3/3 |
| Desert western | 16:9 | 3/3 | 3/3 |
| Overall | 55% raw / ~70% verified | 89% (97% on 16:9) |
¹ Remaining repeats lost to transient network errors, not generation failures. ² Across both grid orientations tested. ³ Raw judge failures later verified as diegetic-text false positives.
Round 4 — vertical showdown (identical prompt, 2 scenarios × 3 repeats):
| Model | Anime chase | Photoreal busker | Overall | Mean latency |
|---|---|---|---|---|
| Seedream 4.5 | 1/3 | 1/3 | 2/6 | ~16s |
| Nano Banana Pro | 1/3 | 0/3 | 1/6 | ~44s |
Round 5 — panel detection on the vertical sheets (15 sheets from rounds 3–4):
| Metric | Blind grid cropping | Detected-panel cropping |
|---|---|---|
| Sheets with every panel usable | ~30% | 27%⁴ |
| Sheets yielding ≥4 usable panels | — | 73% |
| Wrong-count sheets recovered | 0 | 100% (8/8 panels each) |
⁴ Strict all-panels bar with the unguarded prototype; the production detector adds plausibility checks and a uniform-grid fallback. The product-relevant bar is the 4+ row: users select panels, so a sheet with four strong options and two rejects is a working sheet.
A 2×2 vertical probe (6 sheets) mostly held geometry (4 clean cells in 4/6) but lost shot diversity, consistent with the round-1 finding that fewer panels invite repetition.
The through-line of all five rounds: current image models are already excellent cinematographers and inconsistent typesetters. Give them one canvas and one cast, name the layout priors you don't want, verify with your own eyes before trusting a judge — and when the model insists on drawing the grid its own way, measure where it drew the lines instead of insisting on yours.