A new Creative Writing benchmark pits humans against 24 LLMs

475-Prompt Writing Showdown Reveals Only Top AI Models Can Narrowly Outperform Amateur Humans

New AI Creative Writing Benchmark Shows Top LLMs Can Beat Amateur Writers, But Professionals Still Lead

A new creative writing benchmark from Vulsar AI is offering a fresh look at how today’s large language models perform when challenged with long-form storytelling. As AI tools become more common in writing, editing, marketing, and content creation, many have questioned whether human writers could eventually be replaced by advanced language models.

The results suggest a more nuanced answer: the strongest AI models are now capable of outperforming amateur human writers in some creative writing tasks, but professional writers still remain ahead when it comes to depth, consistency, and storytelling skill.

The benchmark tested large language models across 475 writing prompts, focusing on long-form generation rather than short answers or simple text completion. Vulsar AI used a reward model trained on human preference data to estimate which outputs readers would likely prefer.

At the top of the overall leaderboard was GPT 6 Astra, which achieved an 87.8 percent predicted win rate. That placed it slightly above amateur human writers, who scored an 86.6 percent predicted win rate. While the margin is small, it shows that frontier AI models are becoming increasingly competitive in creative writing.

However, the benchmark also makes one point very clear: amateur writers and professional writers are not in the same category. When professional human writers were evaluated separately, they still performed significantly better than AI models and amateur writers. This highlights the gap between producing readable text and crafting polished, emotionally rich, and structurally strong writing.

Other high-ranking AI models also performed well, though they trailed behind GPT 6 Astra. OpenAI’s GPT 5.6 Sol reached a 77.6 percent predicted win rate, while Anthropic’s Claude Fable 5.1 scored 70.3 percent. These results show that only the most advanced AI systems are currently approaching high-level creative writing performance.

The scores dropped sharply among less powerful models. Claude Opus 5, Kimi K3, Grok 4.6, and several others landed in the 50 to 65 percent range. This suggests that even well-known AI models still struggle to consistently match strong human writing across longer creative tasks.

Smaller and less dense models performed much worse. Qwen3.8-27B scored 23.2 percent, DeepSeek V4.1 Flash reached 19.8 percent, and Gemma 4 26B came in at just 10.9 percent. These lower scores point to a major limitation in creative AI writing: generating longer stories requires more than grammar and sentence structure. It demands continuity, pacing, character development, emotional tone, and the ability to maintain coherence over many pages.

The benchmark also compared token output per prompt. Human writers produced the most text on average, with 2,592 tokens per prompt. Claude Fable 5.1 followed with 2,114 tokens, while GPT 6 Astra generated an average of 1,537 tokens. This matters because longer output can reveal weaknesses that shorter samples hide. A model may sound impressive in a few paragraphs but begin to repeat itself, lose track of plot details, or drift in style during multi-chapter writing.

This is where human writers still have a clear advantage. Skilled writers can adapt their voice, adjust pacing, introduce new ideas naturally, and maintain narrative structure over time. Professional writers also bring intuition, lived experience, and creative decision-making that current AI systems still struggle to fully replicate.

The findings do not mean AI is unimportant in creative writing. In fact, the benchmark shows that frontier models are becoming powerful tools for drafting, brainstorming, outlining, and assisting with storytelling. But the results also show that the best human writers still hold an edge in originality, control, and long-form narrative quality.

For now, AI may be able to outperform casual or amateur writing in certain benchmark conditions, but professional creative writing remains a much harder target. The bigger question is how long that advantage will last as large language models continue to improve.