AI video models
We replaced two dancers with AI characters: Wan 3.0, Kling, MiniMax, Omni and Aleph
One dance video, two fictional portraits, five video-editing models. Watch the source and outputs, compare measured costs and resolution, and see where character consistency and moderation broke the workflow.
- video editing
- character consistency
- Wan 3.0
- Kling 3.0 Omni
- MiniMax H3
- video generation costs
Can you take an existing dance video, replace both people with fictional characters, and keep their identities stable as they move and cross each other?
We tried it with five models. Wan 3.0 was our preferred result for this particular scene, and its settled generation charge was $0.975. MiniMax H3 also returned a full-length result, for $2.40. Our earlier Kling 3.0 Omni attempt produced recognizable replacements, but their appearance changed noticeably where two independently generated clips met. Google Gemini Omni 1.1 Flash and Runway Aleph 2 rejected our requests and produced no videos.
This is a practical case study with one source scene, not a general ranking. The inputs and workflows were not identical across every model. Below are the videos, the measured results, and the differences that matter when interpreting them.
The source and the two replacement characters
The source is a vertical dance video: a man begins in the foreground, a woman enters, and the two eventually cross positions. The room, gestures, expressions and camera framing give us concrete things to compare.
Source supplied for the experiment, with its existing TikTok attribution visible. We used the first 14.9 seconds at 576 × 1024, 30 fps. The export outro was trimmed before generation.
Download the original source, including its ending · Download the exact trimmed input
The replacements were these two supplied fictional character portraits:
We asked each model to replace the dancers' faces, hair and upper-body clothing, while keeping the man's dark shorts and the woman's original lower-body garment. We also asked it to preserve the dance, room and lighting, track each identity through crossings, and remove the moving TikTok overlays. The outro removal was preprocessing; it was not a model achievement.
Read the complete common prompt. Reference labels were adapted to each provider's syntax. Aleph used a separate keyframe instruction, and Google's second attempt used a shorter prompt.
Watch the results
These are the returned Wan and MiniMax videos, with their provider audio intact. Kling's player shows our assembled version with the original source audio restored. These are different audio treatments; an audio stream's presence alone does not establish that the original track was preserved.
Wan 3.0: our preferred result in this scene
Wan returned a single 15-second, 720 × 1280 video at 30 fps. The provider reported about 14 minutes 53 seconds of runtime and a settled charge of $0.975.
Our review favored how it retained the portraits' distinctive faces, freckles, hair and fuzzy clothing. In the sampled frames, the two characters remain recognizable after changing positions, without the abrupt reset seen at the join in our Kling version.
It did not follow every instruction. The woman's original green skirt is replaced by a longer sweater-like outfit, and some expressions look distorted. No TikTok overlays were visible in the sampled frames. This is a preferred take, not a claim of flawless editing.
MiniMax H3: a full-length result with different compromises
MiniMax returned 15.08 seconds at 768 × 1344 and 24 fps, from a 15-second request at its 768P tier. Its provider runtime was about 16 minutes 15 seconds, and the settled charge was $2.40.
It preserved the woman's green skirt in the sampled frames, an instruction Wan missed. Both characters acquired the requested hairstyles and sweaters. However, their faces appeared smoother and less faithful to the distinctive portrait textures and proportions. Our reviewer preferred Wan overall. The movement also unfolds differently across the samples: matching a source video's general action is not the same as preserving every pose at the same timestamp. We did not measure frame-by-frame motion error.
No TikTok overlays were visible in the sampled MiniMax frames. Neither this observation nor the visual preference is a repeatability measurement.
Kling 3.0 Omni: why the characters changed halfway through
Our earlier Kling experiment split the source into 7.5-second and 7.4-second inputs. We reused those results for this comparison rather than paying to generate them again. Each returned clip was 7.375 seconds at 720 × 1280 and 24 fps. We adjusted them to their source segment lengths, joined them, and restored the source audio; the assembled player is 14.9 seconds at 30 fps.
Both requests included the exact same two portrait URLs, in the same order, and the exact same prompt. The second request did not include an output frame or any other visual state from the first render.
Around the 7.5-second join, the man's face and hair and the characters' sweater patterns change noticeably. The woman's skirt differs between the two segments too. Reusing portraits helped make the characters recognizable, but did not lock the two independent interpretations together.
That is a weakness of this workflow as well as a result to inspect. It would be misleading to call it proof that Kling cannot maintain identity in a single uninterrupted generation. The two provider runtimes were about 3 minutes 17 seconds and 3 minutes 34 seconds; they were submitted in parallel, so those durations should not be added to describe wall-clock wait.
Raw Kling clip 1 · Raw Kling clip 2
The runs that produced no video
Google Gemini Omni 1.1 Flash: we submitted the first 10 seconds and both
portraits through Google's direct API. The detailed prompt returned HTTP
400 with content_blocked. A second, shorter instruction for the same edit
received the same response. Google described sensitive words in the prompt
but did not identify them. We cannot attribute the failure to a particular
word, portrait, or requested change.
Runway Aleph 2: the Replicate endpoint accepts edited scene keyframes. We independently composed a guide frame at five seconds using the source frame and the two portraits, then submitted it with the 14.9-second video. No Kling output was used to create this guide.
Aleph failed after about 15 seconds. Its outer error was generic, but the
underlying provider log reported SAFETY.INPUT.MULTIMODAL: input media did
not pass moderation. It did not identify whether the video or guide frame
triggered the rejection. There is no Aleph output to rate.
An earlier Seedance 2.0 attempt was also unsuccessful. That provider
specifically returned InputImageSensitiveContentDetected.PrivacyInformation,
flagging an input image as potentially containing a real person. That is a
more specific diagnosis than the Google or Aleph messages. We did not test
all Seedance variants, so this is not a claim about the whole family.
Actual outputs and charges
Costs below belong to the providers actually used. These were not OpenRouter renders: Wan and MiniMax ran through Pika, Kling and Aleph through Replicate, and Omni through Google directly. We are standardizing subsequent experiments on OpenRouter where the required capability exists; that does not change the provenance or billing of these runs.
| Model | Measured output | Provider runtime | Cost evidence | Result |
|---|---|---|---|---|
| Wan 3.0 | 15 s · 720 × 1280 · 30 fps | 14m 53s | $0.975 settled | Preferred take; skirt instruction missed |
| MiniMax H3 | 15.08 s · 768 × 1344 · 24 fps | 16m 15s | $2.40 settled | Full-length take; less faithful facial details |
| Kling 3.0 Omni | Two 7.375 s clips · 720 × 1280 · 24 fps | 3m 17s / 3m 34s | ≈$2.48 historical estimate; $0 new generation spend | Visible change at the join |
| Gemini Omni 1.1 Flash | No output; 10 s input | Rejected in about 19s / 16s | Settled charge not established | Both prompts blocked |
| Aleph 2 | No output; 14.9 s input | Failed in about 15s | Settled charge not established; guide cost also unknown | Input-media moderation |
| Seedance 2.0, earlier attempt | No output | Not included in timed comparison | Settled charge not established | Input image flagged as potentially a real person |
The two successful new video jobs settled at $3.375 combined. This is not the total cost of the whole experiment: it excludes historical Kling spend, any unsettled or unknown failed-attempt charges, guide-frame preparation, subscriptions and taxes. MiniMax's billing response counted 15 input seconds and 15 output seconds.
Duration, resolution and pricing trade-offs
These are endpoint-specific specifications and price references checked on October 5, 2026. They are not promises that every provider exposes the same features. In particular, a model's maximum generated duration can differ from the maximum video it can edit.
| Model / pricing provider | Source-video limit | Output duration | Resolution options | API price basis |
|---|---|---|---|---|
| Wan 3.0 / Pika | 15 s total references; Alibaba documents input + output ≤30 s | 2–30 s within combined limit | 480p, 720p, 1080p | Pika: $0.0325 / $0.065 / $0.13 per output second; our 720p job settled at $0.975 |
| MiniMax H3 / Pika | 15 s total across up to 3 videos | Pika schema: 4–15 s | 768P, 2K | $0.08 / $0.13 per second of both input and output; first 5 image refs included |
| Gemini Omni 1.1 Flash / Google | 10 s uploaded video for editing | 3–10 s generation; edit intended to retain source length | 360p, 720p; 1080p and 4K upscaled | $1.50/M input tokens, $17.50/M video-output tokens; approximately $1.014 for 10 s at 720p, plus inputs |
| Kling 3.0 Omni / Replicate | 3–10 s input for editing | Edit follows source; other generation modes up to 15 s | Tested standard 720p; pro 1080p | Standard $0.168/output second; pro $0.224 without generated audio; historical invoice not retrieved |
| Aleph 2 / Runway list-price reference | 2–30 s; Replicate input below 16 MB | Matches source | Source resolution, up to 1080p | Runway list rate $0.28/s, $0.56 minimum; about $4.17 for 14.9 s, excluding guide preparation; not a verified Replicate charge |
Sources: Pika Wan pricing, Pika H3 pricing, Pika API schema, Alibaba Wan guide, Google Omni specifications, Google pricing, Kling model and pricing, Aleph schema, Runway input limits, and Runway pricing.
What we would change for longer videos
The clearest lesson from Kling is that the same reference portraits do not guarantee the same rendered character across independently generated segments. A longer editing workflow needs to retain both the source motion and the appearance established in earlier output.
We would test sequential segments with shared character references, selected frames from the preceding output, and overlapping source footage so the join can be inspected before assembly. Whether those images act as soft references or strict boundary frames depends on the model and endpoint. This experiment did not test that improved workflow, and it does not establish that such a workflow would eliminate drift.
For this scene, one continuous Wan render avoided our artificial Kling join and gave us the take we preferred. It also missed a wardrobe constraint. Keeping the source and all returned videos visible makes both conclusions possible without reducing the comparison to one score.
Scope and reproducibility
We sampled frames at roughly one-second intervals for visual notes and measured file dimensions, duration and frame rate with ffprobe. The overall preference is subjective and unblinded. There were no repeated successful seeds, no quantitative motion score, and no systematic audio-fidelity test. Higher resolution, different prompts, different providers or another scene could change the outcome.
The comparison also has unequal conditions: Kling used two independent chunks; Omni received only ten seconds; Aleph received a separately generated scene guide; and output resolutions differed. A blocked request is an operational result, not a visual-quality score.
Download the public run record for settings, measured outputs, settled charges and media hashes. The full provider requests remain in our internal experiment archive; credentials and account-specific media URLs are not included in the public record.