§ models · storyline
GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 only in custom test
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3, scoring 38.3 percent in custom tests but only 7.8 percent officially versus Opus 5's 30.2 percent.
OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only through its own API with retained reasoning and context compaction. In the official test environment, the model managed just 7.8 percent.
Opus 5 hit its 30.2 percent without such aids. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness appeared first on The Decoder.
§ sources3 publications · timeline below
- the-decoder.comOpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harnessprimary
- openai.comOpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8% (OpenAI)
- the-decoder.comOpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
§ how this story moved
- primary — The Decoder publishes the launch post.
- The Decoder picks up coverage.
- OpenAI picks up coverage.