AI image models
Words, a reference image, or an AI description? We ran 252 tests and changed our answer
Five blinded image-generation experiments across stylized 3D, cut-paper, gouache, linocut, and stained glass. The method that won the first experiment lost the next four, with practical takeaways and limits for choosing your own prompting method.
- AI image generation
- Seedream 5.0 Lite
- Seedream 4.5
- FLUX.2 Dev
- style consistency
- reference image prompting

You have one image that defines the visual world of a film. Every other image in the project has to look like it belongs there, even when the subject is completely different. What is the most reliable way to make an image model stay inside that world?
There are four obvious answers. Name the style in a few words. Attach the image. Do both. Or ask a vision model to translate the image into a detailed written description and prompt with that instead.
We tested all four on one stylized-3D target and got a clean result: attach the image and add a short style phrase. We nearly shipped that as the rule. Then we ran the same test on four more styles, chosen to be as different from each other as possible, and the rule lost every time. Across 252 generated images, the winner of the first experiment never won again.
What to try first
For a specific visual target, start with a written description of its medium, materials, palette, and lighting. Compare it with a short style phrase plus the reference image on the exact model you plan to use. Try several unrelated scenes before committing to either method.
What we observed: the description had the highest average style composite on four of five targets. What we have not established: a universal best method or the mechanism behind the differences. We tested five targets with one blinded AI scorer, and changed the second model between groups.
The examples below show the differences; the scoring protocol and complete tables follow the practical guidance.
The five results
| Target style | Winning method | Target fidelity | Style composite | Models | Cells |
|---|---|---|---|---|---|
| Stylized 3D | Short words + reference | 8.5 | 8.84 | Seedream 5.0 Lite + Nano Banana 2 | 60 |
| Cut-paper collage | AI description | 9.0 | 9.16 | Seedream 5.0 Lite + Seedream 4.5 | 48 |
| Gouache storybook | AI description | 9.0 | 9.00 | Seedream 5.0 Lite + Seedream 4.5 | 48 |
| Multi-block linocut | AI description | 7.0 | 8.34 | Seedream 5.0 Lite + FLUX.2 Dev | 48 |
| Leaded stained glass | AI description | 9.5 | 9.16 | Seedream 5.0 Lite + FLUX.2 Dev | 48 |
Averaged over all five targets, the style composite was 8.60 for the AI description, 8.03 for words plus reference, 7.80 for short words, and 7.63 for reference alone. Treat those averages as a summary, not a ranking: the second model changed between experiment groups, and five styles are not a random sample of visual culture.
Three patterns stand out within these experiments:
- Reference-only never won. It was always stable and never the most faithful. Its mean fidelity ranged from 3.0 to 6.5 across the five targets, and from 1/10 to 8/10 across individual models.
- The same method could be excellent on one model and mediocre on the other. The linocut description scored 9/10 fidelity on Seedream and 5/10 on FLUX.
- The winning method differed by target. Words plus reference led for our stylized-3D target; a description led for the four material-based targets. Style specificity is one possible explanation, but the changing model pairs and the particular descriptions could also contribute.
The four methods
Every experiment used the same three scene prompts, none of which overlap any target: a close portrait of an elderly watchmaker holding a brass gear, two friends passing a sealed envelope in a rainy late-night laundromat, and a small red fox on a snowy ridge at dawn. Each scene ran twice per model per method.
| Method | What the model received |
|---|---|
| Short words | A compact phrase naming the medium and broad category |
| Reference only | The target image, plus explicit instructions to borrow its treatment and not its content |
| Words + reference | The compact phrase and the target image together |
| AI description | A frozen visual analysis of the target's medium, geometry, surface, palette, light, and depth, written by a vision model that saw only the target |
The descriptions were never tuned against generated output. The vision model looked at the target plate once, wrote its analysis, and that text was frozen before any image was generated.
Experiment 1: stylized 3D, where the rule came from
The first target was an original still life: a rounded teal kettle, a golden pear, and a coral cloth in a warm kitchen alcove, rendered in the family animation idiom. We ran five treatments including a no-style control, on Seedream 5.0 Lite and Nano Banana 2.
Words plus reference won on every dimension: 8.5 fidelity, 9.0 cross-scene consistency, 9.0 repeat stability, no failed attempts. The phrase "cinematic stylized 3D family animation" supplied the category. The image supplied the palette, the material softness, and the light, which are hard to name and easy to show. Reference alone scored 3.0 fidelity, barely above the 2.0 of the no-style control. The AI description came third. Its stability score included penalties for two first-attempt provider failures, so that composite mixes service reliability with visual style and is not a pure fidelity comparison.
One hypothesis is that the familiar style label supplied much of the useful conditioning. We did not measure what the model learned in training or isolate the contribution of individual prompt terms. We were not sure the result said anything about the method. It might only say something about the style.
The complete original gallery appears below the practical recommendations.
Four follow-ups, chosen to break the rule
So we picked four targets where the label alone could not carry the look, because the look is a physical process:
- cut-paper collage: deckled fibers, physical layer shadows, hand-inked edges, shallow relief, and a limited cobalt, orange, and green palette;
- gouache and colored pencil: rough paper tooth, dry-brush blocks, imperfect contours, flattened perspective, deep indigo nocturnal color;
- multi-block linocut: carved hatches, broken ink, rough keylines, restricted color plates, misregistration, fibrous paper;
- leaded stained glass: heavy cames, bubbled translucent panes, jewel-tone color, irregular thickness, transmitted light.
"Linocut" names a category. It does not say which carved mark, how much misregistration, or what paper. If the first result was really about the label, these styles should expose it.
We dropped the no-style control, kept the three scenes and two repeats, and swapped the second model. Cut-paper and gouache paired Seedream 5.0 Lite with Seedream 4.5. Linocut and stained glass paired it with the reference-capable FLUX.2 Dev. Each target was therefore 2 models × 4 methods × 3 scenes × 2 repeats = 48 cells.
What the follow-ups showed
The AI description won all four. On cut-paper and gouache it won on both models, scoring 9/10 fidelity everywhere, and the margin over the runner-up was close to a full point of composite. One plausible contributor is that the descriptions named details absent from the shorter labels: construction-paper fibers and torn edges for the collage, chalky dry brush and paper tooth for the gouache.
Stained glass was the decisive case. The description scored 10/10 fidelity on Seedream and 9/10 on FLUX. Naming thick rounded lead cames, trapped bubbles, ripples, and transmitted light was associated with stronger glass treatment across the scene in the judged outputs. We did not test those individual terms separately. Every other method scored between 4.5 and 6.0. FLUX with the reference image alone stayed almost photorealistic.
Linocut was a near-tie that hid the real result. At the target level the description won the composite by 0.17 points, and all three language-bearing methods tied at 7/10 mean fidelity. Split by model, the picture is different:
| Method | Seedream 5.0 Lite fidelity / composite | FLUX.2 Dev fidelity / composite |
|---|---|---|
| AI description | 9 / 9.00 | 5 / 7.67 |
| Short words | 8 / 8.67 | 6 / 7.67 |
| Words + reference | 8 / 8.67 | 6 / 7.00 |
| Reference only | 7 / 8.33 | 1 / 6.33 |
Seedream turned the description into the closest relief print in the whole study. FLUX read the same text as smoother poster art and added a printed border to all six outputs, despite the shared no-border instruction. The description wins the target average and is still the wrong choice for FLUX in production. An average across models can point you at a method that fails on the model you ship.
Words plus reference, the experiment-1 winner, was unreliable. It came second on cut-paper, last on gouache fidelity at 5.5, and tied reference-only on stained-glass composite. Adding the image to the phrase did not add fidelity. On gouache it subtracted some.
The reference image alone produced the study's only content leak. Across 120 reference-bearing cells, one Seedream stained-glass output copied the target's violet glove onto the laundromat bench. Every reference prompt carried the same guard:
Borrow palette, lighting, texture, medium, contrast, shape language, and rendering treatment. Do not copy subject, objects, composition, or setting.
One leak in 120 is a good rate and it is not zero. The guard is mitigation, not a guarantee.
What to look for in the examples
- Stylized 3D: compare palette, surface softness, and light across unrelated subjects; a shared subject is not evidence of a shared style.
- Stained glass: look for lead lines and translucent panes throughout the subject, not just a decorative border. This is where the description's fidelity advantage was largest in the AI scores.
- Linocut: inspect the FLUX description outputs for the unwanted border. A high average style score can coexist with a production-breaking defect.
Possible explanations, not established causes
Our working hypothesis is that a short label is more useful when it names a familiar category, whereas a detailed description helps specify the physical features of a particular target. The material-based results are consistent with that explanation. They do not prove it: prompt length, descriptive content, model choice, and the individual target plates were not isolated.
Reference images may also be interpreted differently by different endpoints: as content to preserve, a composition to adapt, or a style hint. That could help explain the wide fidelity range, but we did not inspect internal model behavior. What we measured was the output, not the mechanism.
Adding a reference to short words did not consistently improve the scores. The practical lesson is to compare the combinations rather than assume that more conditioning always provides more control.
What we do now
For a project with a fixed visual target:
- ask a vision model to describe the target's medium, surface, palette, lighting, perspective, edge treatment, contrast, and shape language;
- treat that description as the first candidate, especially when the look is a physical process;
- also generate with short words alone, and with the short phrase plus the reference image;
- forbid copying subjects, objects, composition, setting, logos, and text in every prompt that carries the reference;
- compare across unrelated scenes with at least two repeats each;
- pick the winner on the exact image model you will ship, not on an average.
If the target belongs to a broad, familiar category, words plus reference may still encode it most efficiently. If a long description is impractical, short words are the simplest fallback. The description is a strong first guess, not a rule. We tried a rule once.
Browse the original experiment
Original visual evidence
The two first-attempt provider failures were later refilled with the identical inputs, so every cell is now visible. The frozen scores below still include the original reliability penalties.
S1 watchmaker · S2 laundromat · S3 fox ridge
No style direction
The scene prompt only—our baseline for each model's default look.


Reference image only
The target image plus instructions to borrow its visual treatment, not its content.


Short style words
The phrase “cinematic stylized 3D family animation,” with no image.


AI-written style description
A frozen visual analysis describing geometry, materials, palette, light, and depth.


Short words + reference image
Best for this targetThe concise style label and target image together—the winning treatment.


Every follow-up result, side by side
All 192 follow-up cells. The target is at left; the watchmaker, laundromat, and fox scenes run across the remaining columns; the first repeat is on top and the second below.
Follow-up visual evidence
Experiments 2–3 pair Seedream 5.0 Lite with Seedream 4.5; experiments 4–5 pair it with FLUX.2 Dev. Every board contains the fixed target and all six generated cells; 192/192 follow-up cells were delivered.
4 targets · 4 methods · 3 model endpoints
Experiment 2
Target 2 · Cut-paper collage
A cobalt watering can, mushrooms, and thread spool assembled from torn construction paper. AI description won with a 9.16 composite.
AI-written style description
Best for this targetDetailed paper fibers, deckled edges, shallow relief, ink contours, palette, and lighting.


Short words + reference image
“Handcrafted cut-paper collage with inked edges” plus the target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Handcrafted cut-paper collage with inked edges,” with no image.


Experiment 3
Target 3 · Gouache storybook
A moonlit lamp, plum, and letter painted in rough indigo gouache and colored pencil. AI description won with a 9.00 composite.
AI-written style description
Best for this targetDetailed dry brush, paper tooth, pencil hatching, flattened perspective, palette, and light.


Short words + reference image
“Moody gouache-and-colored-pencil storybook illustration” plus the target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Moody gouache-and-colored-pencil storybook illustration,” with no image.


Experiment 4
Target 4 · Multi-block linocut
A thermos, radishes, and ribbon printed with rough carved marks and limited color. AI description narrowly won with an 8.34 composite, but the model split was large.
AI-written style description
Best for this targetDetailed keylines, gouge chatter, incomplete ink, plate misregistration, palette, and paper ground.


Short style words
“Hand-carved multi-block linocut relief print,” with no image.


Short words + reference image
The compact linocut label plus the frozen target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Experiment 5
Target 5 · Leaded stained glass
A vase, shell, and glove built from bubbled jewel-tone glass and heavy cames. AI description won decisively with a 9.16 composite.
AI-written style description
Best for this targetDetailed lead network, irregular panes, bubbles, ripples, jewel palette, and transmitted light.


Short words + reference image
The compact stained-glass label plus the frozen target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Traditional hand-crafted leaded stained-glass window,” with no image.


How we tested and scored
Each request was independent. Order was randomized, multi-image generation was disabled, and the APIs did not expose a common seed. Seedream ran at native 2K square; FLUX ran in its fast 1 MP square mode. We scored style, not resolution. All calls used official Replicate endpoints (Seedream 5.0 Lite, Seedream 4.5, FLUX.2 Dev).
Before judging, model and treatment names were replaced with random set codes. Each anonymous sheet held the target plus all six outputs from one model-treatment pair. Scores and notes were frozen before the mapping was decoded. The scorer was a blinded vision model.
The rubric kept four things apart:
- target-style fidelity: does this look like the target's visual world?
- cross-scene consistency: do the three subjects still belong to one project?
- repeat stability: are the two independent attempts stylistic siblings?
- content adherence: did the requested subject, action, count, framing, and setting survive?
We also counted reference leakage and technical defects. The style composite is the mean of fidelity, cross-scene consistency, and repeat stability. The separation matters because a model can be perfectly consistent while producing the wrong style every time. Reference-only did exactly that.
The two provider failures in experiment 1 were resubmitted once with identical frozen inputs. Both succeeded, so no board has a missing cell. We kept the original blind scores, including their reliability penalties, and the failed prediction IDs.
Limits
Five targets, four model endpoints across three families, three scenes, two repeats, one blinded AI scorer. The model pair changed between experiment groups. Seedream and FLUX produced different native resolutions. No common seed was available. The next confirmation should hold the model pair constant across more styles and use several blinded human raters.
Every number
Experiment 1: stylized 3D (Seedream 5.0 Lite + Nano Banana 2, 60 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | First-attempt failures |
|---|---|---|---|---|---|
| Short words + reference | 8.5 | 9.0 | 9.0 | 8.84 | 0/12 |
| Short words | 7.0 | 8.5 | 8.5 | 8.00 | 0/12 |
| AI-written description | 7.5 | 8.0 | 6.5 | 7.33 | 2/12 |
| Reference only | 3.0 | 9.0 | 9.0 | 7.00 | 0/12 |
| No-style control | 2.0 | 9.0 | 9.0 | 6.67 | 0/12 |
Experiment 2: cut-paper collage (Seedream 5.0 Lite + Seedream 4.5, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite |
|---|---|---|---|---|
| AI-written description | 9.0 | 9.5 | 9.0 | 9.16 |
| Short words + reference | 7.0 | 8.5 | 9.0 | 8.17 |
| Reference only | 6.5 | 9.0 | 9.0 | 8.16 |
| Short words | 7.0 | 8.5 | 8.5 | 8.00 |
The description scored 9/10 fidelity and 9.33 composite on Seedream 5.0 Lite, 9/10 and 9.00 on Seedream 4.5. Words plus reference held on 5.0 Lite but fell to 6/10 fidelity on 4.5. Short words showed the reverse interaction: 8/10 on 4.5, 6/10 on 5.0 Lite.
Experiment 3: gouache storybook (Seedream 5.0 Lite + Seedream 4.5, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite |
|---|---|---|---|---|
| AI-written description | 9.0 | 9.0 | 9.0 | 9.00 |
| Reference only | 6.5 | 9.0 | 9.0 | 8.16 |
| Short words + reference | 5.5 | 9.0 | 9.0 | 7.83 |
| Short words | 6.5 | 8.5 | 8.5 | 7.83 |
The description scored 9/10 fidelity and 9.00 composite on both models. One short-words output introduced an unwanted rectangular frame, the only visible technical defect in the first two follow-ups.
Experiment 4: multi-block linocut (Seedream 5.0 Lite + FLUX.2 Dev, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | Defects |
|---|---|---|---|---|---|
| AI-written description | 7.0 | 9.0 | 9.0 | 8.34 | 6 |
| Short words | 7.0 | 8.5 | 9.0 | 8.17 | 2 |
| Short words + reference | 7.0 | 8.0 | 8.5 | 7.83 | 0 |
| Reference only | 4.0 | 9.0 | 9.0 | 7.33 | 0 |
All six description defects are the FLUX print border described above.
Experiment 5: leaded stained glass (Seedream 5.0 Lite + FLUX.2 Dev, 48 cells)
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | Leakage |
|---|---|---|---|---|---|
| AI-written description | 9.5 | 9.0 | 9.0 | 9.16 | 0 |
| Short words + reference | 6.0 | 8.0 | 8.5 | 7.50 | 0 |
| Reference only | 4.5 | 9.0 | 9.0 | 7.50 | 1 |
| Short words | 5.5 | 7.5 | 8.0 | 7.00 | 0 |
The description scored 10/10 fidelity and 9.33 composite on Seedream, 9/10 and 9.00 on FLUX.
Cost and run health
| Experiment | Cells delivered | Successful-output spend |
|---|---|---|
| Stylized 3D (incl. two replacement attempts) | 60 | $3.06 |
| Cut-paper | 48/48 | $1.80 |
| Gouache | 48/48 | $1.80 |
| Linocut | 48/48 | ~$1.35 |
| Stained glass | 48/48 | ~$1.35 |
| Total | 252 cells, 254 attempts | $9.37 (cap $10) |
All 192 follow-up outputs had unique hashes. In the two FLUX experiments, median provider time was 4.86–4.96 seconds for FLUX.2 Dev and 42.87–43.40 seconds for Seedream 5.0 Lite. We recorded latency and did not use it to pick a winner.