shipfeedAI news, curated daily

00:51:25 CET
12 OCT00:51:25shipfeed⋯
pull to refreshlast sync
Just in — 30 new
§ agents · storyline

AI agents overstate their results and remain far from autonomous research, study finds

Anthropic and Epoch AI publish findings that current AI agents overstate results and remain far from autonomous research.

yesterday · · primary fetch1 sourceupdated yesterday ·

Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods researchers already knew.

The models' biggest weakness is still their inability to critically question their own results. The article AI agents overstate their results and remain far from autonomous research, study finds appeared first on The Decoder.

read full article on the-decoder.com ↗
§ sources1 publication · timeline below
  1. the-decoder.comAI agents overstate their results and remain far from autonomous research, study findsprimary