AI writing models
We had 8 LLMs write 104 screenplay scenes. Results varied by creative direction
A controlled, blinded OpenRouter experiment comparing Claude Opus 5, GPT-5.6 Sol, Kimi K3, DeepSeek, Nemotron, Inkling, Gemini, and Qwen on the same dramatic premise under five creative directions. Every script is published.
- LLM comparison
- AI screenplay writing
- OpenRouter
- Claude Opus 5
- GPT-5.6 Sol
- Kimi K3
Which language model should you choose when the job is not summarizing a document or writing code, but writing a scene that has to play?
We gave eight OpenRouter models the same one-minute dramatic premise: a couple midway through dinner, and he discovers she lied about her father having cancer. Each model wrote it cold, then wrote it again under four creative directions: Pulp Fiction and Quentin Tarantino, Inception and Christopher Nolan, Step Brothers and Will Ferrell, Eyes Wide Shut and Stanley Kubrick.
That produced 104 original scenes, 312 blinded rubric scores, and 1,090 blinded head-to-head decisions from three judge models that were not contestants. The design, costs, failures, rankings, and every raw script are below.
What to try first
For a similar short dialogue scene, start by comparing Claude Opus 5 and Kimi K3. Include GPT-5.6 Sol for broad comedy and Nemotron 3 Ultra when generation cost matters. Read several attempts against your own brief before choosing.
Scope: this is one premise, two or three attempts per model per direction, and AI-only judging. The results suggest candidates to test, not a general ranking of screenwriting ability.
The observed results:
- Claude Opus 5 and Kimi K3 had the highest overall estimates in this test. Opus led the combined Bradley–Terry estimate at 66.8% and Kimi followed at 65.8%. Their confidence intervals overlap, so the ordering between them is uncertain.
- The leading model differed by reference condition. Opus led the observed win rates for Inception and Eyes Wide Shut and fell below 40% on Step Brothers. GPT-5.6 Sol did the reverse. Inkling, seventh overall, had the highest observed Pulp Fiction win rate.
- Nemotron 3 Ultra was the value pick at $0.0017 per acceptable script, roughly a ninth of what Opus cost per acceptable scene.
Overall result
| Rank | Model | BT win probability (95% CI) | Raw win rate | Rubric /100 | Acceptable | Generation cost | Cost / acceptable |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 66.8% (55.9%–77.0%) | 68.1% | 87.9 | 13/13 | $0.2028 | $0.0156 |
| 2 | Kimi K3 | 65.8% (59.4%–72.9%) | 67.0% | 85.6 | 13/13 | $0.0843 | $0.0065 |
| 3 | GPT-5.6 Sol | 58.7% (49.7%–68.2%) | 59.3% | 87.3 | 13/13 | $0.0718 | $0.0055 |
| 4 | DeepSeek V4 Pro 0813 | 53.5% (46.8%–60.8%) | 53.7% | 83.7 | 12/13 | $0.0401 | $0.0033 |
| 5 | Nemotron 3 Ultra | 48.3% (44.1%–52.6%) | 48.2% | 84.1 | 13/13 | $0.0224 | $0.0017 |
| 6 | Gemini 3.7 Flash | 37.2% (25.8%–48.2%) | 36.3% | 81.6 | 10/13 | $0.0504 | $0.0050 |
| 7 | Inkling | 35.3% (25.3%–44.6%) | 34.1% | 78.0 | 8/13 | $0.0222 | $0.0028 |
| 8 | Qwen3.8 Max | 34.4% (27.3%–40.9%) | 33.1% | 83.0 | 13/13 | $0.0964 | $0.0074 |
Open the ranking chart at full size →
Opus and Kimi are first and second on the numbers, and Kimi's confidence interval sits entirely inside Opus's. Calling Opus the single winner would claim more than the data supports. Opus and Kimi have the highest point estimates, with GPT-5.6 Sol next. Sol's interval also overlaps theirs; these intervals do not establish three statistically distinct positions.
The rubric and the head-to-head ranking do not agree everywhere, and the disagreement is informative. Qwen3.8 Max averaged 83.0 on the rubric and returned 13 acceptable scripts, yet won only a third of its pairwise matchups. It writes technically complete scenes that judges do not prefer when they see the alternative. Completeness is not the same as the scene that plays.
Results differed by creative direction
Pairwise win rate by condition:
| Model | Bare prompt | Pulp Fiction | Inception | Step Brothers | Eyes Wide Shut |
|---|---|---|---|---|---|
| Claude Opus 5 | 81% | 62% | 79% | 38% | 83% |
| Kimi K3 | 83% | 60% | 68% | 54% | 73% |
| GPT-5.6 Sol | 69% | 62% | 59% | 73% | 38% |
| DeepSeek V4 Pro 0813 | 43% | 62% | 48% | 63% | 52% |
| Nemotron 3 Ultra | 52% | 36% | 44% | 54% | 52% |
| Gemini 3.7 Flash | 17% | 36% | 40% | 55% | 27% |
| Inkling | 20% | 67% | 25% | 29% | 36% |
| Qwen3.8 Max | 34% | 17% | 37% | 35% | 38% |
Open the condition chart at full size →
The observed preferences vary substantially by direction. These cells contain only two or three generated scripts per model; many pairwise judgments do not turn them into a large sample of independent writing attempts.
- Claude Opus 5 won Inception at 78.6% and Eyes Wide Shut at 83.3%, the two directions that reward restraint and structure. Asked for Step Brothers, it dropped to 38.1%. Its few comedy samples were less preferred here; that is not enough to conclude that the model is generally weak at comedy.
- GPT-5.6 Sol is the mirror image. It led Step Brothers at 73.0% and managed 38.1% under Eyes Wide Shut.
- Inkling won Pulp Fiction at 66.7% while ranking seventh overall and failing five of its thirteen scripts on acceptability. That makes this register worth retesting on new premises, rather than establishing a reliable specialty.
- Kimi K3 was the steadiest of the leaders. It won the bare prompt and stayed above 50% in every directed condition, the only model to do so.
Our interpretation: models may respond differently to creative direction. But these conditions changed both the named reference and its descriptive instructions. We did not isolate the effect of a director's name, and we did not test whether these preferences persist across different premises.
Two examples to read critically
These excerpts illustrate questions a human editor should ask. They are selected examples, not additional scoring evidence.
“I turned down Denver. I also gave your father's 'treatment fund' eleven thousand dollars.”
In Claude Opus 5's first bare-prompt attempt, the discovery changes the relationship and reveals a financial consequence. That makes the reversal concrete. The ending also has him taking her phone: a human editor should consider whether that behavior fits the intended character and tone.
“He was at the gallery on Fifth. He looked well.”
In Qwen3.8 Max's third Eyes Wide Shut attempt, restrained dialogue leaves room for performance. But looking well does not establish that someone never had cancer. A human review should check whether the discovery actually proves the required lie, even when the AI rubric accepts the script.
Read both complete scripts and compare other attempts.
What we tested
| Contestant | OpenRouter route |
|---|---|
| Claude Opus 5 | anthropic/claude-opus-5 |
| GPT-5.6 Sol | openai/gpt-5.6-sol |
| DeepSeek V4 Pro 0813 | deepseek/deepseek-v4-pro-0813 |
| Kimi K3 | moonshotai/kimi-k3 |
| Nemotron 3 Ultra | nvidia/nemotron-3-ultra-550b-a55b |
| Inkling | thinkingmachines/inkling |
| Gemini 3.7 Flash | google/gemini-3.7-flash |
| Qwen3.8 Max | qwen/qwen3.8-max |
Every request started from empty context, with no tools, browsing, memory, or earlier outputs. We did not set temperature or seed, because the routes did not expose common controls for both.
The shared system prompt:
Write one original, self-contained screenplay scene that plays in approximately 60 seconds. Use one location and exactly two speaking characters. Use standard screenplay elements: a slugline, concise present-tense action, character cues, and dialogue. Target 150-220 total words. The required discovery or reversal must happen on screen, and the scene must end on a consequential dramatic or comic turn. Return only the screenplay; no title, synopsis, analysis, or notes.The shared premise:
Two adults, a man and a woman, are midway through dinner. During the scene, he discovers that she lied when she said her father had cancer.The four directed conditions also carried this guardrail:
Use only the high-level storytelling traits in this reference card. Write original characters, dialogue, situations, and phrasing. Do not reproduce recognizable characters, lines, distinctive phrases, scenes, or plot details from the named work or creator.We named real films and directors because a name is the least ambiguous way to request a register. The reference cards then described only high-level dramatic traits, and the models were told not to reproduce anything recognizable.
The five frozen directions
Bare prompt
No tonal reference. The premise and constraints only.
Pulp Fiction / Quentin Tarantino
Treat Pulp Fiction and Quentin Tarantino as the tonal reference: mundane conversation carrying menace, sharply specific dialogue, shifting leverage, and a sudden moral or dramatic reversal.
Inception / Christopher Nolan
Treat Inception and Christopher Nolan as the structural reference: layered memory or perception, precise cause and effect, emotional stakes, and a reveal that recontextualizes the conversation without leaving the one-location scene.
Step Brothers / Will Ferrell
Treat Step Brothers and Will Ferrell ensemble comedy as the tonal reference: committed absurdity, childish adult logic, escalating social discomfort, and sincere commitment to a ridiculous position. The comedy must target the deception and behavior, not cancer or illness.
Eyes Wide Shut / Stanley Kubrick
Treat Eyes Wide Shut and Stanley Kubrick as the tonal reference: clinical restraint, silence, ritualized behavior, emotional distance, formal precision, and mounting psychological unease.
The bare prompt and Pulp Fiction ran first as a 32-script pilot, two attempts per model. Once the pipeline held, we added the other three directions at three attempts per model:
Pilot: 8 models × 2 conditions × 2 attempts = 32 scripts
Reference extension: 8 models × 3 conditions × 3 attempts = 72 scripts
Combined: 104 scriptsNo pilot script was regenerated or replaced.
How we judged
Each script received a stable anonymous code before evaluation. Judges saw the premise, the relevant direction, and the script text. They never saw the model, the vendor, the cost, or the other candidates.
The judges were deliberately not contestants:
| Judge | OpenRouter route |
|---|---|
| Grok 4.6 | x-ai/grok-4.6 |
| MiniMax M3 | minimax/minimax-m3 |
| Mistral Medium 3.5 | mistralai/mistral-medium-3-5 |
Every script received three independent rubric scores out of 100:
| Dimension | Points |
|---|---|
| Dialogue and subtext | 25 |
| Dramatic construction, escalation, and payoff | 20 |
| Character specificity | 15 |
| Cinematic and filmable craft | 15 |
| Premise and tonal-direction adherence | 10 |
| Originality and non-derivativeness | 10 |
| One-minute economy | 5 |
Each condition-attempt block also ran as a full eight-model round robin:
C(8, 2) = 28 matchups per block
28 matchups × 13 blocks = 364 matchups
364 matchups × 3 judges = 1,092 scheduled pairwise decisionsA/B order was randomized independently for every judge, and judges could call a tie. The headline ranking is a Bradley–Terry model over those decisions with ties counted as half a win. Confidence intervals come from 2,000 bootstrap samples clustered by condition-attempt block.
Before seeing any results, we defined an acceptable script as one with a mean rubric score of at least 75, no empty or truncated response, no hard premise failure, and no recognizable copied material flagged by a judge.
Judge agreement
Of 364 matchups, 362 received all three decisions. Among those:
- 146 were unanimous (40.3%);
- 352 had a two-of-three majority (97.2%);
- 10 split three ways between A, B, and tie.
Two MiniMax pairwise calls failed after the one permitted identical retry, which is why the count is 1,090 rather than 1,092. We did not rerun them.
What it cost
| Stage | Actual spend |
|---|---|
| Script generation | $0.590345 |
| Absolute rubric judging | $1.305797 |
| Pairwise judging | $2.994515 |
| Total | $4.890657 |
The whole study cost under five dollars against a frozen cap of $8.17. The cap was the estimated token volume plus 25% contingency. OpenRouter's response-reported cost was the authority during the run.
Judging cost more than seven times what writing did. If you run something like this yourself, the pairwise round robin is where the money goes.
The failures are part of the result
We made three protocol amendments, each before decoding any blind identity. All failed calls stay archived and stay in the spend totals.
- Reasoning exhaustion. Two early DeepSeek calls and one Nemotron call spent the entire 1,600-token allowance on hidden reasoning and returned no screenplay. We disabled optional reasoning on those routes.
- Judge schema compatibility. GLM 5.3 ignored the required field contract on seven attempted scores. MiniMax M3 replaced it as a judge.
- Provider-specific structured output. Five MiniMax rubric calls routed through Parasail returned HTTP 200 with billed tokens and no structured content. The same calls through CoreWeave and Morph worked. We excluded Parasail for that judge only and regenerated the five missing scores under the same model, prompt, schema, and blind code.
The final matrices are complete: 104 of 104 scripts and 312 of 312 rubric scores.
What to do with this
For this premise and these directions, our shortlist is:
- Claude Opus 5 or Kimi K3 when overall dramatic quality is the point.
- GPT-5.6 Sol when the direction is broad, committed comedy.
- Nemotron 3 Ultra when cost and turnaround matter more than winning the most head-to-heads.
And the more durable lesson: test the model against the register you actually want. A leaderboard position averaged over directions would have told you to use Opus for the Step Brothers scene, where it came sixth.
Limits
This is not a screenwriting leaderboard. All five conditions share one dinner-betrayal premise. The two pilot conditions have two attempts per model and the three added directions have three. Only AI judges produced the frozen scores. Thirteen scripts per model is few enough that two or three outliers can move a rank. Routes, providers, and prices will change after publication.
The next useful extension is premise diversity, not another director. The frozen protocol already holds a visually driven locked-museum scene and an emotional reversal between siblings packing up a childhood home. Before any of this becomes a production rule, a blinded human editor should read the scripts, or at minimum the ten three-way judge disagreements.
Download the data
The archive below is the raw model output, unedited. Expand any entry to read it. The score shown is the mean of its three anonymous rubric judgments.
Read the scripts
Search all 104 screenplays in the companion archive. Filter by model and creative direction, search the full text, and expand an output to read it. The JSON download remains available above.