OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings the-decoder.com
43 stories · 7d·6 sources covering·30 active storylines
What this is
Evals are the benchmarks and tests that measure how AI models perform on reasoning, coding, and safety tasks. shipfeed tracks new benchmark releases and notable evaluation results.
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings the-decoder.com
Artificial Intelligence Achieves Its First Perfect Score at the International Mathematical Olympiad (2026 Shanghai Competition) XenoSpectrum
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.
Moonshot AI's Kimi K3 Beats Claude and GPT Rivals to Top a Coding Benchmark Startup Fortune
China’s Kimi K3 Is Out—And Beats Claude Fable and GPT 5.6 Sol on Key Benchmarks Decrypt
An OpenAI model crushed top human programmers at a world coding competition understandingai.org
OpenAI Launches GPT-5.6 AI Model Family teleSUR English
Nimbus builds production AI systems — internal tools, customer agents, retrieval pipelines — combining humans and AI end-to-end. From scoped pilot to production in 4–8 weeks.
Anthropic Launches Claude Science AI Workbench For Scientists Pulse 2.0
Artificial Analysis: GLM-5.2 is the leading open weights model on Artificial Analysis' Intelligence Index, scoring 51, only behind Fable 5's 60, Opus 4.8's 56, and GPT-5.5's 55 — Z ai's GLM-5.2 is the new leading…
Ram Iyer / TechCrunch: Microsoft releases ASSERT, an open-source framework that lets developers generate and run AI behavior tests using natural-language descriptions — AI researchers and labs have advanced by…
Madison Mills / Axios: Anthropic says it expects Mythos-class models to be available to all customers “in the coming weeks” following the development of stronger safeguards — Anthropic released Claude…
Researchers at Carnegie Mellon University built a new benchmark that measures how far AI agents can go when exploiting real vulnerabilities in Google's V8 engine. Mythos leads GPT-5.5 by a wide margin but costs twelve…
AI Security Institute: Mythos Preview is the first AI model to complete both of AISI's cyber ranges, which measure models' cyberattack capabilities; GPT-5.5 solved only one of them — In February 2026, we internally…
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
OpenAI’s GPT-5.6 Tests Show Prompt-Injection Gains and Agent Risks TechRepublic
Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent to do and what the agent thought.You…
Microsoft In-House Cyber Model Beats Anthropic and OpenAI on Security Benchmark at Half Cost Tech Times
Luz Ding / Bloomberg: Alibaba says its 2.4T-parameter Qwen3.8-Max tops Kimi K3 on some benchmarks, and plans to release the open weights of Qwen3.8-Max and Qwen3.8-27B next week — Alibaba Group Holding Ltd…
DeepSeek’s Cheap Model Just Beat Its Own Flagship on Nine Benchmarks Times Tabloid
Nimbus builds production AI systems — internal tools, customer agents, retrieval pipelines — combining humans and AI end-to-end. From scoped pilot to production in 4–8 weeks.
Opus 5 Tops New AI Leaderboard: What Developers Need to Know SitePoint
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a…
Top AIs invent same fake PyPl and npm package names InfoWorld