§ evals · storyline
New benchmark finds only frontier LLMs beat amateur writers
Benchmark pitting 24 LLMs against human writers across 475 prompts finds only frontier models outperform amateurs.
New Benchmark Pits Human Writers Against 24 LLMs Across 475 Prompts, Results Show That Only Frontier Models Barely Edge Past Amateur Individuals, Showcasing A Major Skill Gap Wccftech
§ sources1 publication · timeline below