§ agents · storyline

UK's AI Security Institute finds benchmarks underestimate agent

UK AI Security Institute finds standard benchmarks systematically underestimate AI agent capabilities by limiting compute budgets.

yesterday · 18:14:44 · primary fetch1 sourceupdated yesterday · 18:14:44

In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks, success rates jumped about 25 percent when the token budget was increased tenfold. Newer models benefit the most.

Depending on the token budget, actual progress at the frontier is about 60 percent steeper than previous measurements suggested, according to AISI. The article UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do appeared first on The Decoder.

read full article on the-decoder.com ↗

§ sources1 publication · timeline below

the-decoder.comUK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually doprimary18:14:44