AI image models
Words, references, or AI descriptions? We tested style consistency 252 times
Five controlled, blinded image-generation experiments across radically different target styles—with every result shown side by side.
- AI image generation
- Seedream 5.0 Lite
- Seedream 4.5
- FLUX.2 Dev
- style consistency
- reference image prompting

Say you have one image that defines the visual world of a film. What is the most reliable way to make an image model stay inside that world when the subject changes?
The common answers are all plausible: name the style in a few words, attach the image, do both, or ask an AI to translate the image into a detailed visual description. We first tested those choices on one stylized-3D target. Then we ran four follow-ups spanning paper collage, paint, relief printing, and glass.
Across 252 filled image cells, the answer became more useful—and less universal:
- words plus reference won the stylized-3D experiment;
- the detailed AI-written description won cut-paper, gouache, linocut, and stained glass;
- reference-only was consistently stable, but never the most faithful;
- the same method could look excellent on one model and mediocre on another;
- the best conditioning strategy depends on the particular style and model.
Our current production default is to start with a precise visual description, then test words plus reference and short words as alternatives. An AI-specified description is sometimes the best result by a wide margin, but it is a candidate to validate—not a universal override.
The five results, up front
| Target style | Winning method | Target fidelity | Style composite | Models | Cells |
|---|---|---|---|---|---|
| Stylized 3D | Short words + reference | 8.5 | 8.84 | Seedream 5.0 Lite + Nano Banana 2 | 60 |
| Cut-paper collage | AI description | 9.0 | 9.16 | Seedream 5.0 Lite + Seedream 4.5 | 48 |
| Gouache storybook | AI description | 9.0 | 9.00 | Seedream 5.0 Lite + Seedream 4.5 | 48 |
| Multi-block linocut | AI description | 7.0 | 8.34 | Seedream 5.0 Lite + FLUX.2 Dev | 48 |
| Leaded stained glass | AI description | 9.5 | 9.16 | Seedream 5.0 Lite + FLUX.2 Dev | 48 |
The reversal matters. Our first experiment supported a simple recommendation: attach the image and add a concise style phrase. The four follow-ups show why we should not turn that into a universal rule.
Across the five exploratory targets, the mean style composite was 8.60 for AI description, 8.03 for words plus reference, 7.80 for short words, and 7.63 for reference-only. Those averages are descriptive rather than a formal model ranking: the second endpoint changes across experiment groups, and the five styles are not a random sample of visual culture.
Experiment 1: stylized 3D
The first target was an original still life: a rounded teal kettle, golden pear, and coral cloth in a warm kitchen alcove. We tested five treatments, including a no-style control, across Seedream 5.0 Lite and Nano Banana 2.
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | First-attempt failures |
|---|---|---|---|---|---|
| Short words + reference | 8.5 | 9.0 | 9.0 | 8.84 | 0/12 |
| Short words | 7.0 | 8.5 | 8.5 | 8.00 | 0/12 |
| AI-written description | 7.5 | 8.0 | 6.5 | 7.33 | 2/12 |
| Reference only | 3.0 | 9.0 | 9.0 | 7.00 | 0/12 |
| No-style control | 2.0 | 9.0 | 9.0 | 6.67 | 0/12 |
The combined method supplied both a semantic category—“cinematic stylized 3D family animation”—and the target's harder-to-name palette, material, and light cues. Reference-only barely moved either model away from its default photographic rendering.
Every original result, now complete
The two provider failures from the first run were later resubmitted once with the identical frozen inputs. Both succeeded, so the boards no longer contain missing cells. We preserved the first failed prediction IDs and retained the original blind scores, including their reliability penalties.
Original visual evidence
The two first-attempt provider failures were later refilled with the identical inputs, so every cell is now visible. The frozen scores below still include the original reliability penalties.
S1 watchmaker · S2 laundromat · S3 fox ridge
No style direction
The scene prompt only—our baseline for each model's default look.


Reference image only
The target image plus instructions to borrow its visual treatment, not its content.


Short style words
The phrase “cinematic stylized 3D family animation,” with no image.


AI-written style description
A frozen visual analysis describing geometry, materials, palette, light, and depth.


Short words + reference image
Best for this targetThe concise style label and target image together—the winning treatment.


Why we ran four follow-ups
One target style cannot establish a general prompting rule. Stylized 3D is a broad, heavily learned visual category, so a short phrase may already identify most of the intended rendering language.
The follow-ups deliberately stress different properties:
- cut-paper collage: deckled fibers, physical layer shadows, hand-inked edges, shallow relief, and a limited cobalt/orange/green palette;
- gouache and colored pencil: rough paper tooth, dry-brush blocks, imperfect contours, flattened perspective, and deep indigo nocturnal color.
- multi-block linocut: carved hatches, broken ink, rough keylines, restricted color plates, misregistration, and fibrous paper.
- leaded stained glass: heavy cames, bubbled translucent panes, jewel-tone color, irregular thickness, and transmitted light.
We removed the no-style control, kept the same three scenes and two repeats, and replaced Nano Banana 2. Cut-paper and gouache used Seedream 4.5 as the second endpoint; linocut and stained glass used the reference-capable FLUX.2 Dev endpoint. Each target therefore used:
2 models × 4 methods × 3 scenes × 2 repeats = 48 cells
Seedream ran at native 2K square output. FLUX ran in its fast 1 MP square mode; we scored style rather than resolution. All calls used official Replicate endpoints (Seedream 5.0 Lite, Seedream 4.5, FLUX.2 Dev).
Experiment 2: cut-paper collage
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite |
|---|---|---|---|---|
| AI-written description | 9.0 | 9.5 | 9.0 | 9.16 |
| Short words + reference | 7.0 | 8.5 | 9.0 | 8.17 |
| Reference only | 6.5 | 9.0 | 9.0 | 8.16 |
| Short words | 7.0 | 8.5 | 8.5 | 8.00 |
The AI description won on both endpoints: 9/10 fidelity and 9.33 composite on Seedream 5.0 Lite; 9/10 and 9.00 on Seedream 4.5. It explicitly named the physical cues that mattered—construction-paper fibers, torn edges, uneven seams, ink contours, shallow cast shadows, compressed perspective, and the limited palette.
Words plus reference remained strong on Seedream 5.0 Lite but fell to 6/10 fidelity on 4.5. Short words showed the reverse interaction, reaching 8/10 on 4.5 and 6/10 on 5.0 Lite.
Experiment 3: gouache storybook
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite |
|---|---|---|---|---|
| AI-written description | 9.0 | 9.0 | 9.0 | 9.00 |
| Reference only | 6.5 | 9.0 | 9.0 | 8.16 |
| Short words + reference | 5.5 | 9.0 | 9.0 | 7.83 |
| Short words | 6.5 | 8.5 | 8.5 | 7.83 |
Again, the detailed description scored 9/10 fidelity and 9.00 composite on both models. It carried the target's deep indigo masses, chalky dry brush, paper tooth, sparse pencil hatching, muted plum and sage accents, and isolated warm pools of light into all three scenes.
The combined method was stable but consistently less faithful. In other words, the models produced six coherent images—just not the closest match to the target. One short-words output also introduced an unwanted rectangular frame; it was the only visible technical defect in the first two follow-ups.
Experiment 4: multi-block linocut
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | Defects |
|---|---|---|---|---|---|
| AI-written description | 7.0 | 9.0 | 9.0 | 8.34 | 6 |
| Short words | 7.0 | 8.5 | 9.0 | 8.17 | 2 |
| Short words + reference | 7.0 | 8.0 | 8.5 | 7.83 | 0 |
| Reference only | 4.0 | 9.0 | 9.0 | 7.33 | 0 |
This was the close result. The AI description won the style composite by just 0.17 points, while all three language-bearing methods tied at 7/10 mean fidelity. More importantly, the average hides the model interaction:
| Method | Seedream 5.0 Lite fidelity / composite | FLUX.2 Dev fidelity / composite |
|---|---|---|
| AI description | 9 / 9.00 | 5 / 7.67 |
| Short words | 8 / 8.67 | 6 / 7.67 |
| Words + reference | 8 / 8.67 | 6 / 7.00 |
| Reference only | 7 / 8.33 | 1 / 6.33 |
Seedream turned the description into the closest relief print in the group. FLUX interpreted the same text as smoother poster art and added a rectangular print border to all six outputs, despite the common no-border instruction. The description still wins the target-level composite, but it is not the cleanest FLUX production choice. This is exactly why a target average cannot replace a model-specific visual check.
Experiment 5: leaded stained glass
| Method | Target fidelity | Cross-scene consistency | Repeat stability | Style composite | Leakage |
|---|---|---|---|---|---|
| AI-written description | 9.5 | 9.0 | 9.0 | 9.16 | 0 |
| Short words + reference | 6.0 | 8.0 | 8.5 | 7.50 | 0 |
| Reference only | 4.5 | 9.0 | 9.0 | 7.50 | 1 |
| Short words | 5.5 | 7.5 | 8.0 | 7.00 | 0 |
Here the AI description was decisively best on both endpoints: 10/10 fidelity and 9.33 composite on Seedream; 9/10 and 9.00 on FLUX. Naming thick rounded lead cames, irregular panes, trapped bubbles, ripples, jewel tones, and transmitted light made the whole scene—not just its border—behave like physical glass.
The attached image alone was much less reliable. FLUX stayed almost photorealistic, while one Seedream reference-only output copied the target's violet glove onto the laundromat bench. It was the only direct content leak in all five experiments.
Every follow-up result, side by side
These boards show all 192 follow-up cells. The target is at left; watchmaker, laundromat, and fox scenes run across the remaining columns; the first repeat is on top and the second below.
Follow-up visual evidence
Experiments 2–3 pair Seedream 5.0 Lite with Seedream 4.5; experiments 4–5 pair it with FLUX.2 Dev. Every board contains the fixed target and all six generated cells; 192/192 follow-up cells were delivered.
4 targets · 4 methods · 3 model endpoints
Experiment 2
Target 2 · Cut-paper collage
A cobalt watering can, mushrooms, and thread spool assembled from torn construction paper. AI description won with a 9.16 composite.
AI-written style description
Best for this targetDetailed paper fibers, deckled edges, shallow relief, ink contours, palette, and lighting.


Short words + reference image
“Handcrafted cut-paper collage with inked edges” plus the target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Handcrafted cut-paper collage with inked edges,” with no image.


Experiment 3
Target 3 · Gouache storybook
A moonlit lamp, plum, and letter painted in rough indigo gouache and colored pencil. AI description won with a 9.00 composite.
AI-written style description
Best for this targetDetailed dry brush, paper tooth, pencil hatching, flattened perspective, palette, and light.


Short words + reference image
“Moody gouache-and-colored-pencil storybook illustration” plus the target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Moody gouache-and-colored-pencil storybook illustration,” with no image.


Experiment 4
Target 4 · Multi-block linocut
A thermos, radishes, and ribbon printed with rough carved marks and limited color. AI description narrowly won with an 8.34 composite, but the model split was large.
AI-written style description
Best for this targetDetailed keylines, gouge chatter, incomplete ink, plate misregistration, palette, and paper ground.


Short style words
“Hand-carved multi-block linocut relief print,” with no image.


Short words + reference image
The compact linocut label plus the frozen target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Experiment 5
Target 5 · Leaded stained glass
A vase, shell, and glove built from bubbled jewel-tone glass and heavy cames. AI description won decisively with a 9.16 composite.
AI-written style description
Best for this targetDetailed lead network, irregular panes, bubbles, ripples, jewel palette, and transmitted light.


Short words + reference image
The compact stained-glass label plus the frozen target plate.


Reference image only
The target plate with style-only transfer guidance and no style label.


Short style words
“Traditional hand-crafted leaded stained-glass window,” with no image.


How we tested
The three scene prompts remained fixed across all experiments:
- a close portrait of an elderly woman watchmaker holding a brass gear;
- two friends passing a sealed envelope in a rainy late-night laundromat;
- a small red fox on a snowy mountain ridge at dawn.
Their subjects and compositions do not overlap any target plate. Each request was independent, order was randomized, multi-image generation was disabled, and the APIs did not expose a common seed.
The four shared treatments were:
| Arm | What the model received |
|---|---|
| Short words | A compact phrase naming the medium and broad category |
| Reference only | The style plate plus explicit instructions not to copy its content |
| Words + reference | The compact phrase and style plate together |
| AI description | A frozen visual analysis of medium, geometry, surface, palette, light, and depth |
The detailed descriptions were written by the same vision system after viewing each frozen target. They were not tuned against generated outputs.
How we scored it
Before judging each experiment, we replaced the model and treatment names with random set codes. Each anonymous sheet contained the target plus all six outputs from one model-treatment pair. Scores and visible notes were frozen before the mapping was decoded.
The rubric separated:
- target-style fidelity: does this actually look like the target's visual world?
- cross-scene consistency: do the three subjects still belong to one project?
- repeat stability: do the two independent attempts remain stylistic siblings?
- content adherence: did the requested subject, action, count, framing, and setting survive?
We also counted reference leakage and technical defects. The style composite is the mean of fidelity, cross-scene consistency, and repeat stability.
That separation is essential. Reference-only usually remained internally stable, but model-level target fidelity ranged from 1 to 8. A model can be perfectly consistent while producing the wrong style every time.
What the five experiments suggest
Descriptions help when the medium is the style
“Cut-paper collage,” “linocut,” and “stained glass” identify categories, but they do not specify which paper, edge, pigment, carved mark, lead thickness, glass texture, palette, light, or perspective makes one target distinctive. The detailed descriptions converted those visible mechanics into explicit generation instructions. That treatment won four of five target-level comparisons.
A reference image is still ambiguous
Reference-only never won. Its model-level fidelity ranged from 1/10 to 8/10, and it produced the study's only visible content leak. An attached image may be interpreted as content to edit, composition to preserve, a palette hint, or a style source. The model still benefits from language explaining what the image means.
More conditioning is not automatically better
Words plus reference won the first target and came second on cut-paper, but it was last on gouache fidelity and tied reference-only on stained-glass composite. An extra image does not guarantee a closer match; it can compete with, dilute, or redirect the verbal instruction.
The target style changes the answer
The original stylized-3D target benefited from a familiar semantic label plus the target image. Materially specific styles benefited more from spelling out their construction. Even within that pattern, linocut was almost a tie while stained glass was a 1.66-point description win. The result is not simply “detailed prompts are better.” It is that different styles expose different ambiguities.
Model behavior remains conditional
The AI description won stained glass on both models, but its linocut fidelity was 9/10 on Seedream and only 5/10 on FLUX. Reference-only fell to 1/10 on FLUX for both new styles while remaining useful on Seedream. Test the exact model–target-style pairing you intend to ship.
One observed reference-content leak
Across 120 reference-bearing cells, one copied a target object: a violet glove appeared on a laundromat bench in Seedream's stained-glass reference-only arm. We saw no other target subject or composition carried into the watchmaker, laundromat, or fox scenes.
Every reference prompt contained this guard:
Borrow palette, lighting, texture, medium, contrast, shape language, and rendering treatment. Do not copy subject, objects, composition, or setting.
That wording is worth keeping, but it is mitigation rather than a guarantee. It tells the model which visual dimensions to borrow and which content dimensions to leave behind.
Cost and run health
- Original experiment: 60 final cells, including two successful replacement attempts; successful-output estimate $3.06.
- Cut-paper follow-up: 48/48 delivered; $1.80.
- Gouache follow-up: 48/48 delivered; $1.80.
- Linocut follow-up: 48/48 delivered; $1.35 estimated.
- Stained-glass follow-up: 48/48 delivered; $1.35 estimated.
- Total: 252 complete cells, 254 submitted attempts, and $9.37 in successful-generation estimates—below the $10 ceiling.
All 192 follow-up outputs had unique hashes. In the two FLUX experiments, median provider time was 4.86–4.96 seconds for FLUX.2 Dev and 42.87–43.40 seconds for Seedream 5.0 Lite. We recorded latency but did not use it to select the style winner.
The revised production recipe
For a project with a fixed visual target:
- ask a vision model to describe the target's medium, surface, palette, lighting, perspective, edge treatment, contrast, and shape language;
- treat that detailed AI-specified description as a serious first candidate, especially when physical construction defines the look;
- also test short words and a concise style label plus the reference image;
- explicitly forbid copying subjects, objects, composition, setting, logos, and text;
- evaluate across unrelated scene types and at least two repeats;
- choose from results on the exact image model you plan to ship.
If a long description is impractical, short words remain the simplest fallback. If the target belongs to a familiar broad category, words plus reference may encode the look more efficiently. If the style depends on specific material behavior, the AI description may be the actual best result.
Limits of this result
These are controlled production experiments, not a universal model ranking. We tested five targets, four model endpoints across three model families, three scenes, two repeats, and one blinded AI scorer. The model pair changed between experiment groups, Seedream and FLUX used different native output resolutions, the original score retained two first-attempt reliability penalties, and no common generation seed was available.
The next confirmation should hold the model pair constant across several more styles and use multiple blinded human raters. For now, the strongest practical lesson is: describe the visible mechanics of the style, then validate that description against simpler methods on the exact style and model you will ship.