shipfeedAI news, curated daily

23:50:26 CET
13 AUG23:50:26shipfeed
pull to refreshlast sync
Just in — 30 new
§ models · storyline

GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 only in custom test

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3, scoring 38.3 percent in custom tests but only 7.8 percent officially versus Opus 5's 30.2 percent.

Jul 30 · · primary fetch2 sourcesupdated Jul 30 ·

OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only through its own API with retained reasoning and context compaction. In the official test environment, the model managed just 7.8 percent.

Opus 5 hit its 30.2 percent without such aids. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness appeared first on The Decoder.

read full article on the-decoder.com
§ sources3 publications · timeline below
  1. the-decoder.comOpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harnessprimary
  2. openai.comOpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8% (OpenAI)
  3. the-decoder.comOpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

§ how this story moved

  1. primaryThe Decoder publishes the launch post.
  2. The Decoder picks up coverage.
  3. OpenAI picks up coverage.