AI video models

We replaced two dancers with AI characters: Wan 3.0, Kling, MiniMax, Omni and Aleph

One dance video, two fictional portraits, five video-editing models. Watch the source and outputs, compare measured costs and resolution, and see where character consistency and moderation broke the workflow.

By Sundream
  • video editing
  • character consistency
  • Wan 3.0
  • Kling 3.0 Omni
  • MiniMax H3
  • video generation costs

Can you take an existing dance video, replace both people with fictional characters, and keep their identities stable as they move and cross each other?

We tried it with five models. Wan 3.0 was our preferred result for this particular scene, and its settled generation charge was $0.975. MiniMax H3 also returned a full-length result, for $2.40. Our earlier Kling 3.0 Omni attempt produced recognizable replacements, but their appearance changed noticeably where two independently generated clips met. Google Gemini Omni 1.1 Flash and Runway Aleph 2 rejected our requests and produced no videos.

This is a practical case study with one source scene, not a general ranking. The inputs and workflows were not identical across every model. Below are the videos, the measured results, and the differences that matter when interpreting them.

The source and the two replacement characters

The source is a vertical dance video: a man begins in the foreground, a woman enters, and the two eventually cross positions. The room, gestures, expressions and camera framing give us concrete things to compare.

Source supplied for the experiment, with its existing TikTok attribution visible. We used the first 14.9 seconds at 576 × 1024, 30 fps. The export outro was trimmed before generation.

Download the original source, including its ending · Download the exact trimmed input

The replacements were these two supplied fictional character portraits:

Male character reference: freckled face, thin mustache, tall sculpted dark hair and a mustard, navy and cream sweater

Female character reference: freckled face, dark updo with a colorful wool ornament, gold jewelry and a cream, orange and navy sweater

We asked each model to replace the dancers' faces, hair and upper-body clothing, while keeping the man's dark shorts and the woman's original lower-body garment. We also asked it to preserve the dance, room and lighting, track each identity through crossings, and remove the moving TikTok overlays. The outro removal was preprocessing; it was not a model achievement.

Read the complete common prompt. Reference labels were adapted to each provider's syntax. Aleph used a separate keyframe instruction, and Google's second attempt used a shorter prompt.

Watch the results

These are the returned Wan and MiniMax videos, with their provider audio intact. Kling's player shows our assembled version with the original source audio restored. These are different audio treatments; an audio stream's presence alone does not establish that the original track was preserved.

Wan 3.0: our preferred result in this scene

Wan returned a single 15-second, 720 × 1280 video at 30 fps. The provider reported about 14 minutes 53 seconds of runtime and a settled charge of $0.975.

Our review favored how it retained the portraits' distinctive faces, freckles, hair and fuzzy clothing. In the sampled frames, the two characters remain recognizable after changing positions, without the abrupt reset seen at the join in our Kling version.

It did not follow every instruction. The woman's original green skirt is replaced by a longer sweater-like outfit, and some expressions look distorted. No TikTok overlays were visible in the sampled frames. This is a preferred take, not a claim of flawless editing.

MiniMax H3: a full-length result with different compromises

MiniMax returned 15.08 seconds at 768 × 1344 and 24 fps, from a 15-second request at its 768P tier. Its provider runtime was about 16 minutes 15 seconds, and the settled charge was $2.40.

It preserved the woman's green skirt in the sampled frames, an instruction Wan missed. Both characters acquired the requested hairstyles and sweaters. However, their faces appeared smoother and less faithful to the distinctive portrait textures and proportions. Our reviewer preferred Wan overall. The movement also unfolds differently across the samples: matching a source video's general action is not the same as preserving every pose at the same timestamp. We did not measure frame-by-frame motion error.

No TikTok overlays were visible in the sampled MiniMax frames. Neither this observation nor the visual preference is a repeatability measurement.

Kling 3.0 Omni: why the characters changed halfway through

Our earlier Kling experiment split the source into 7.5-second and 7.4-second inputs. We reused those results for this comparison rather than paying to generate them again. Each returned clip was 7.375 seconds at 720 × 1280 and 24 fps. We adjusted them to their source segment lengths, joined them, and restored the source audio; the assembled player is 14.9 seconds at 30 fps.

Both requests included the exact same two portrait URLs, in the same order, and the exact same prompt. The second request did not include an output frame or any other visual state from the first render.

Around the 7.5-second join, the man's face and hair and the characters' sweater patterns change noticeably. The woman's skirt differs between the two segments too. Reusing portraits helped make the characters recognizable, but did not lock the two independent interpretations together.

That is a weakness of this workflow as well as a result to inspect. It would be misleading to call it proof that Kling cannot maintain identity in a single uninterrupted generation. The two provider runtimes were about 3 minutes 17 seconds and 3 minutes 34 seconds; they were submitted in parallel, so those durations should not be added to describe wall-clock wait.

Raw Kling clip 1 · Raw Kling clip 2

The runs that produced no video

Google Gemini Omni 1.1 Flash: we submitted the first 10 seconds and both portraits through Google's direct API. The detailed prompt returned HTTP 400 with content_blocked. A second, shorter instruction for the same edit received the same response. Google described sensitive words in the prompt but did not identify them. We cannot attribute the failure to a particular word, portrait, or requested change.

Runway Aleph 2: the Replicate endpoint accepts edited scene keyframes. We independently composed a guide frame at five seconds using the source frame and the two portraits, then submitted it with the 14.9-second video. No Kling output was used to create this guide.

Aleph guide frame showing both fictional characters in the source room at the five-second pose

Aleph failed after about 15 seconds. Its outer error was generic, but the underlying provider log reported SAFETY.INPUT.MULTIMODAL: input media did not pass moderation. It did not identify whether the video or guide frame triggered the rejection. There is no Aleph output to rate.

An earlier Seedance 2.0 attempt was also unsuccessful. That provider specifically returned InputImageSensitiveContentDetected.PrivacyInformation, flagging an input image as potentially containing a real person. That is a more specific diagnosis than the Google or Aleph messages. We did not test all Seedance variants, so this is not a claim about the whole family.

Actual outputs and charges

Costs below belong to the providers actually used. These were not OpenRouter renders: Wan and MiniMax ran through Pika, Kling and Aleph through Replicate, and Omni through Google directly. We are standardizing subsequent experiments on OpenRouter where the required capability exists; that does not change the provenance or billing of these runs.

ModelMeasured outputProvider runtimeCost evidenceResult
Wan 3.015 s · 720 × 1280 · 30 fps14m 53s$0.975 settledPreferred take; skirt instruction missed
MiniMax H315.08 s · 768 × 1344 · 24 fps16m 15s$2.40 settledFull-length take; less faithful facial details
Kling 3.0 OmniTwo 7.375 s clips · 720 × 1280 · 24 fps3m 17s / 3m 34s≈$2.48 historical estimate; $0 new generation spendVisible change at the join
Gemini Omni 1.1 FlashNo output; 10 s inputRejected in about 19s / 16sSettled charge not establishedBoth prompts blocked
Aleph 2No output; 14.9 s inputFailed in about 15sSettled charge not established; guide cost also unknownInput-media moderation
Seedance 2.0, earlier attemptNo outputNot included in timed comparisonSettled charge not establishedInput image flagged as potentially a real person

The two successful new video jobs settled at $3.375 combined. This is not the total cost of the whole experiment: it excludes historical Kling spend, any unsettled or unknown failed-attempt charges, guide-frame preparation, subscriptions and taxes. MiniMax's billing response counted 15 input seconds and 15 output seconds.

Duration, resolution and pricing trade-offs

These are endpoint-specific specifications and price references checked on October 5, 2026. They are not promises that every provider exposes the same features. In particular, a model's maximum generated duration can differ from the maximum video it can edit.

Model / pricing providerSource-video limitOutput durationResolution optionsAPI price basis
Wan 3.0 / Pika15 s total references; Alibaba documents input + output ≤30 s2–30 s within combined limit480p, 720p, 1080pPika: $0.0325 / $0.065 / $0.13 per output second; our 720p job settled at $0.975
MiniMax H3 / Pika15 s total across up to 3 videosPika schema: 4–15 s768P, 2K$0.08 / $0.13 per second of both input and output; first 5 image refs included
Gemini Omni 1.1 Flash / Google10 s uploaded video for editing3–10 s generation; edit intended to retain source length360p, 720p; 1080p and 4K upscaled$1.50/M input tokens, $17.50/M video-output tokens; approximately $1.014 for 10 s at 720p, plus inputs
Kling 3.0 Omni / Replicate3–10 s input for editingEdit follows source; other generation modes up to 15 sTested standard 720p; pro 1080pStandard $0.168/output second; pro $0.224 without generated audio; historical invoice not retrieved
Aleph 2 / Runway list-price reference2–30 s; Replicate input below 16 MBMatches sourceSource resolution, up to 1080pRunway list rate $0.28/s, $0.56 minimum; about $4.17 for 14.9 s, excluding guide preparation; not a verified Replicate charge

Sources: Pika Wan pricing, Pika H3 pricing, Pika API schema, Alibaba Wan guide, Google Omni specifications, Google pricing, Kling model and pricing, Aleph schema, Runway input limits, and Runway pricing.

What we would change for longer videos

The clearest lesson from Kling is that the same reference portraits do not guarantee the same rendered character across independently generated segments. A longer editing workflow needs to retain both the source motion and the appearance established in earlier output.

We would test sequential segments with shared character references, selected frames from the preceding output, and overlapping source footage so the join can be inspected before assembly. Whether those images act as soft references or strict boundary frames depends on the model and endpoint. This experiment did not test that improved workflow, and it does not establish that such a workflow would eliminate drift.

For this scene, one continuous Wan render avoided our artificial Kling join and gave us the take we preferred. It also missed a wardrobe constraint. Keeping the source and all returned videos visible makes both conclusions possible without reducing the comparison to one score.

Scope and reproducibility

We sampled frames at roughly one-second intervals for visual notes and measured file dimensions, duration and frame rate with ffprobe. The overall preference is subjective and unblinded. There were no repeated successful seeds, no quantitative motion score, and no systematic audio-fidelity test. Higher resolution, different prompts, different providers or another scene could change the outcome.

The comparison also has unequal conditions: Kling used two independent chunks; Omni received only ten seconds; Aleph received a separately generated scene guide; and output resolutions differed. A blocked request is an operational result, not a visual-quality score.

Download the public run record for settings, measured outputs, settled charges and media hashes. The full provider requests remain in our internal experiment archive; credentials and account-specific media URLs are not included in the public record.