AI image models
Words, a reference image, or an AI description? We ran 252 tests and changed our answer
Five blinded image-generation experiments across stylized 3D, cut-paper, gouache, linocut, and stained glass. The method that won the first experiment lost the next four, and the reason is the useful part.
- AI image generation
- Seedream 5.0 Lite
- Seedream 4.5
- FLUX.2 Dev
- style consistency
- reference image prompting

You have one image that defines the visual world of a film. Every other image in the project has to look like it belongs there, even when the subject is completely different. What is the most reliable way to make an image model stay inside that world?
There are four obvious answers. Name the style in a few words. Attach the image. Do both. Or ask a vision model to translate the image into a detailed written description and prompt with that instead.
We tested all four on one stylized-3D target and got a clean result: attach the image and add a short style phrase. We nearly shipped that as the rule. Then we ran the same test on four more styles, chosen to be as different from each other as possible, and the rule lost every time. Across 252 generated images, the winner of the first experiment never won again.
This post is about why, and what we do now instead.
The five results
| Target style | Winning method | Target fidelity | Style composite | Models | Cells |
|---|---|---|---|---|---|
| Stylized 3D | Short words + reference | 8.5 | 8.84 | Seedream 5.0 Lite + Nano Banana 2 | 60 |
| Cut-paper collage | AI description | 9.0 | 9.16 | Seedream 5.0 Lite + Seedream 4.5 | 48 |
| Gouache storybook | AI description | 9.0 | 9.00 | Seedream 5.0 Lite + Seedream 4.5 | 48 |
| Multi-block linocut | AI description | 7.0 | 8.34 | Seedream 5.0 Lite + FLUX.2 Dev | 48 |
| Leaded stained glass | AI description | 9.5 | 9.16 | Seedream 5.0 Lite + FLUX.2 Dev | 48 |
Averaged over all five targets, the style composite was 8.60 for the AI description, 8.03 for words plus reference, 7.80 for short words, and 7.63 for reference alone. Treat those averages as a summary, not a ranking: the second model changed between experiment groups, and five styles are not a random sample of visual culture.
Three things held across every experiment, and they matter more than the averages:
- Reference-only never won. It was always stable and never the most faithful. Its mean fidelity ranged from 3.0 to 6.5 across the five targets, and from 1/10 to 8/10 across individual models.
- The same method could be excellent on one model and mediocre on the other. The linocut description scored 9/10 fidelity on Seedream and 5/10 on FLUX.
- The style decided the winner. A broad, heavily learned category like stylized 3D responds to a label plus an image. A style defined by physical construction, like glass or carved ink, responds to having that construction spelled out.
The four methods
Every experiment used the same three scene prompts, none of which overlap any target: a close portrait of an elderly watchmaker holding a brass gear, two friends passing a sealed envelope in a rainy late-night laundromat, and a small red fox on a snowy ridge at dawn. Each scene ran twice per model per method.
| Method | What the model received |
|---|---|
| Short words | A compact phrase naming the medium and broad category |
| Reference only | The target image, plus explicit instructions to borrow its treatment and not its content |
| Words + reference | The compact phrase and the target image together |
| AI description | A frozen visual analysis of the target's medium, geometry, surface, palette, light, and depth, written by a vision model that saw only the target |
The descriptions were never tuned against generated output. The vision model looked at the target plate once, wrote its analysis, and that text was frozen before any image was generated.
Experiment 1: stylized 3D, where the rule came from
The first target was an original still life: a rounded teal kettle, a golden pear, and a coral cloth in a warm kitchen alcove, rendered in the family animation idiom. We ran five treatments including a no-style control, on Seedream 5.0 Lite and Nano Banana 2.
Words plus reference won on every dimension: 8.5 fidelity, 9.0 cross-scene consistency, 9.0 repeat stability, no failed attempts. The phrase "cinematic stylized 3D family animation" supplied the category. The image supplied the palette, the material softness, and the light, which are hard to name and easy to show. Reference alone scored 3.0 fidelity, barely above the 2.0 of the no-style control. The AI description came third, with two first-attempt provider failures dragging its stability score down.
That is a tidy story. It is also exactly the story you would expect if the style label is doing most of the work, because "stylized 3D family animation" is a broad, heavily learned visual category. We were not sure the result said anything about the method. It might only say something about the style.
Original visual evidence
The two first-attempt provider failures were later refilled with the identical inputs, so every cell is now visible. The frozen scores below still include the original reliability penalties.
S1 watchmaker · S2 laundromat · S3 fox ridge
No style direction
The scene prompt only—our baseline for each model's default look.


Reference image only
The target image plus instructions to borrow its visual treatment, not its content.


Short style words
The phrase “cinematic stylized 3D family animation,” with no image.


AI-written style description
A frozen visual analysis describing geometry, materials, palette, light, and depth.


Short words + reference image
Best for this targetThe concise style label and target image together—the winning treatment.


Four follow-ups, chosen to break the rule
So we picked four targets where the label alone could not carry the look, because the look is a physical process:
- cut-paper collage: deckled fibers, physical layer shadows, hand-inked edges, shallow relief, and a limited cobalt, orange, and green palette;
- gouache and colored pencil: rough paper tooth, dry-brush blocks, imperfect contours, flattened perspective, deep indigo nocturnal color;
- multi-block linocut: carved hatches, broken ink, rough keylines, restricted color plates, misregistration, fibrous paper;
- leaded stained glass: heavy cames, bubbled translucent panes, jewel-tone color, irregular thickness, transmitted light.
"Linocut" names a category. It does not say which carved mark, how much misregistration, or what paper. If the first result was really about the label, these styles should expose it.
We dropped the no-style control, kept the three scenes and two repeats, and swapped the second model. Cut-paper and gouache paired Seedream 5.0 Lite with Seedream 4.5. Linocut and stained glass paired it with the reference-capable FLUX.2 Dev. Each target was therefore 2 models × 4 methods × 3 scenes × 2 repeats = 48 cells.
What the follow-ups showed
The AI description won all four. On cut-paper and gouache it won on both models, scoring 9/10 fidelity everywhere, and the margin over the runner-up was close to a full point of composite. The descriptions named exactly the things the labels leave out: construction-paper fibers and torn edges for the collage, chalky dry brush and paper tooth for the gouache.
Stained glass was the decisive case. The description scored 10/10 fidelity on Seedream and 9/10 on FLUX. Naming thick rounded lead cames, trapped bubbles, ripples, and transmitted light made the whole scene behave like glass, not just its border. Every other method scored between 4.5 and 6.0. FLUX with the reference image alone stayed almost photorealistic.
Linocut was a near-tie that hid the real result. At the target level the description won the composite by 0.17 points, and all three language-bearing methods tied at 7/10 mean fidelity. Split by model, the picture is different:
| Method | Seedream 5.0 Lite fidelity / composite | FLUX.2 Dev fidelity / composite |
|---|---|---|
| AI description | 9 / 9.00 | 5 / 7.67 |
| Short words | 8 / 8.67 | 6 / 7.67 |
| Words + reference | 8 / 8.67 | 6 / 7.00 |
| Reference only | 7 / 8.33 | 1 / 6.33 |
Seedream turned the description into the closest relief print in the whole study. FLUX read the same text as smoother poster art and added a printed border to all six outputs, despite the shared no-border instruction. The description wins the target average and is still the wrong choice for FLUX in production. An average across models can point you at a method that fails on the model you ship.
Words plus reference, the experiment-1 winner, was unreliable. It came second on cut-paper, last on gouache fidelity at 5.5, and tied reference-only on stained-glass composite. Adding the image to the phrase did not add fidelity. On gouache it subtracted some.
The reference image alone produced the study's only content leak. Across 120 reference-bearing cells, one Seedream stained-glass output copied the target's violet glove onto the laundromat bench. Every reference prompt carried the same guard:
Borrow palette, lighting, texture, medium, contrast, shape language, and rendering treatment. Do not copy subject, objects, composition, or setting.
One leak in 120 is a good rate and it is not zero. The guard is mitigation, not a guarantee.
Every follow-up result, side by side
All 192 follow-up cells. The target is at left; the watchmaker, laundromat, and fox scenes run across the remaining columns; the first repeat is on top and the second below.
Follow-up visual evidence
Experiments 2–3 pair Seedream 5.0 Lite with Seedream 4.5; experiments 4–5 pair it with FLUX.2 Dev. Every board contains the fixed target and all six generated cells; 192/192 follow-up cells were delivered.
4 targets · 4 methods · 3 model endpoints
Experiment 2
Target 2 · Cut-paper collage
A cobalt watering can, mushrooms, and thread spool assembled from torn construction paper. AI description won with a 9.16 composite.
AI-written style description
Best for this targetDetailed paper fibers, deckled edges, shallow relief, ink contours, palette, and lighting.


Short words + reference image
“Handcrafted cut-paper collage with inked edges” plus the target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Handcrafted cut-paper collage with inked edges,” with no image.


Experiment 3
Target 3 · Gouache storybook
A moonlit lamp, plum, and letter painted in rough indigo gouache and colored pencil. AI description won with a 9.00 composite.
AI-written style description
Best for this targetDetailed dry brush, paper tooth, pencil hatching, flattened perspective, palette, and light.


Short words + reference image
“Moody gouache-and-colored-pencil storybook illustration” plus the target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Moody gouache-and-colored-pencil storybook illustration,” with no image.


Experiment 4
Target 4 · Multi-block linocut
A thermos, radishes, and ribbon printed with rough carved marks and limited color. AI description narrowly won with an 8.34 composite, but the model split was large.
AI-written style description
Best for this targetDetailed keylines, gouge chatter, incomplete ink, plate misregistration, palette, and paper ground.


Short style words
“Hand-carved multi-block linocut relief print,” with no image.


Short words + reference image
The compact linocut label plus the frozen target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Experiment 5
Target 5 · Leaded stained glass
A vase, shell, and glove built from bubbled jewel-tone glass and heavy cames. AI description won decisively with a 9.16 composite.
AI-written style description
Best for this targetDetailed lead network, irregular panes, bubbles, ripples, jewel palette, and transmitted light.


Short words + reference image
The compact stained-glass label plus the frozen target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Traditional hand-crafted leaded stained-glass window,” with no image.


Why the answer moved
A short label works when the style is a category the model already knows intimately. "Stylized 3D family animation" is that. The label recalls a whole rendering pipeline, and the reference image fine-tunes palette and light on top of it.
A label fails when the style is a physical process with a dozen visible consequences. "Stained glass" recalls a look. It does not specify lead thickness, pane irregularity, or how light should pass through the scene rather than fall on it. The AI description turns those visible mechanics into instructions, which is why it won four of five and won by the most where the material was most specific.
The reference image alone is ambiguous in a way that language is not. A model can read an attached image as content to edit, a composition to preserve, a palette hint, or a style source. Reference-only fidelity ranged from 1/10 to 8/10 across models and targets for exactly that reason. Text telling the model what the image means is what makes it useful.
And more conditioning is not more control. Stacking an image on a phrase can compete with the phrase, dilute it, or redirect it. Words plus reference looked like the safe maximal choice after experiment 1 and turned out to be the least predictable of the four.
What we do now
For a project with a fixed visual target:
- ask a vision model to describe the target's medium, surface, palette, lighting, perspective, edge treatment, contrast, and shape language;
- treat that description as the first candidate, especially when the look is a physical process;
- also generate with short words alone, and with the short phrase plus the reference image;
- forbid copying subjects, objects, composition, setting, logos, and text in every prompt that carries the reference;
- compare across unrelated scenes with at least two repeats each;
- pick the winner on the exact image model you will ship, not on an average.
If the target belongs to a broad, familiar category, words plus reference may still encode it most efficiently. If a long description is impractical, short words are the simplest fallback. The description is a strong first guess, not a rule. We tried a rule once.
How we tested and scored
Each request was independent. Order was randomized, multi-image generation was disabled, and the APIs did not expose a common seed. Seedream ran at native 2K square; FLUX ran in its fast 1 MP square mode. We scored style, not resolution. All calls used official Replicate endpoints (Seedream 5.0 Lite, Seedream 4.5, FLUX.2 Dev).
Before judging, model and treatment names were replaced with random set codes. Each anonymous sheet held the target plus all six outputs from one model-treatment pair. Scores and notes were frozen before the mapping was decoded. The scorer was a blinded vision model.
The rubric kept four things apart:
- target-style fidelity: does this look like the target's visual world?
- cross-scene consistency: do the three subjects still belong to one project?
- repeat stability: are the two independent attempts stylistic siblings?
- content adherence: did the requested subject, action, count, framing, and setting survive?
We also counted reference leakage and technical defects. The style composite is the mean of fidelity, cross-scene consistency, and repeat stability. The separation matters because a model can be perfectly consistent while producing the wrong style every time. Reference-only did exactly that.
The two provider failures in experiment 1 were resubmitted once with identical frozen inputs. Both succeeded, so no board has a missing cell. We kept the original blind scores, including their reliability penalties, and the failed prediction IDs.
Limits
Five targets, four model endpoints across three families, three scenes, two repeats, one blinded AI scorer. The model pair changed between experiment groups. Seedream and FLUX produced different native resolutions. No common seed was available. The next confirmation should hold the model pair constant across more styles and use several blinded human raters.
Every number
Experiment 1: stylized 3D (Seedream 5.0 Lite + Nano Banana 2, 60 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | First-attempt failures |
|---|---|---|---|---|---|
| Short words + reference | 8.5 | 9.0 | 9.0 | 8.84 | 0/12 |
| Short words | 7.0 | 8.5 | 8.5 | 8.00 | 0/12 |
| AI-written description | 7.5 | 8.0 | 6.5 | 7.33 | 2/12 |
| Reference only | 3.0 | 9.0 | 9.0 | 7.00 | 0/12 |
| No-style control | 2.0 | 9.0 | 9.0 | 6.67 | 0/12 |
Experiment 2: cut-paper collage (Seedream 5.0 Lite + Seedream 4.5, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite |
|---|---|---|---|---|
| AI-written description | 9.0 | 9.5 | 9.0 | 9.16 |
| Short words + reference | 7.0 | 8.5 | 9.0 | 8.17 |
| Reference only | 6.5 | 9.0 | 9.0 | 8.16 |
| Short words | 7.0 | 8.5 | 8.5 | 8.00 |
The description scored 9/10 fidelity and 9.33 composite on Seedream 5.0 Lite, 9/10 and 9.00 on Seedream 4.5. Words plus reference held on 5.0 Lite but fell to 6/10 fidelity on 4.5. Short words showed the reverse interaction: 8/10 on 4.5, 6/10 on 5.0 Lite.
Experiment 3: gouache storybook (Seedream 5.0 Lite + Seedream 4.5, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite |
|---|---|---|---|---|
| AI-written description | 9.0 | 9.0 | 9.0 | 9.00 |
| Reference only | 6.5 | 9.0 | 9.0 | 8.16 |
| Short words + reference | 5.5 | 9.0 | 9.0 | 7.83 |
| Short words | 6.5 | 8.5 | 8.5 | 7.83 |
The description scored 9/10 fidelity and 9.00 composite on both models. One short-words output introduced an unwanted rectangular frame, the only visible technical defect in the first two follow-ups.
Experiment 4: multi-block linocut (Seedream 5.0 Lite + FLUX.2 Dev, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | Defects |
|---|---|---|---|---|---|
| AI-written description | 7.0 | 9.0 | 9.0 | 8.34 | 6 |
| Short words | 7.0 | 8.5 | 9.0 | 8.17 | 2 |
| Short words + reference | 7.0 | 8.0 | 8.5 | 7.83 | 0 |
| Reference only | 4.0 | 9.0 | 9.0 | 7.33 | 0 |
All six description defects are the FLUX print border described above.
Experiment 5: leaded stained glass (Seedream 5.0 Lite + FLUX.2 Dev, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | Leakage |
|---|---|---|---|---|---|
| AI-written description | 9.5 | 9.0 | 9.0 | 9.16 | 0 |
| Short words + reference | 6.0 | 8.0 | 8.5 | 7.50 | 0 |
| Reference only | 4.5 | 9.0 | 9.0 | 7.50 | 1 |
| Short words | 5.5 | 7.5 | 8.0 | 7.00 | 0 |
The description scored 10/10 fidelity and 9.33 composite on Seedream, 9/10 and 9.00 on FLUX.
Cost and run health
| Experiment | Cells delivered | Successful-output spend |
|---|---|---|
| Stylized 3D (incl. two replacement attempts) | 60 | $3.06 |
| Cut-paper | 48/48 | $1.80 |
| Gouache | 48/48 | $1.80 |
| Linocut | 48/48 | ~$1.35 |
| Stained glass | 48/48 | ~$1.35 |
| Total | 252 cells, 254 attempts | $9.37 (cap $10) |
All 192 follow-up outputs had unique hashes. In the two FLUX experiments, median provider time was 4.86–4.96 seconds for FLUX.2 Dev and 42.87–43.40 seconds for Seedream 5.0 Lite. We recorded latency and did not use it to pick a winner.