shipfeedAI news, curated daily

03:07:53 CET
28 JUL03:07:53shipfeed
pull to refreshlast sync

Custom · research — shipfeed

About this track

Custom filter — share the URL to share the view.

Custom · research50 storylines

Sunday, July 26, 2026’s edition
Together AI — Blog
GPT · 1 source

Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.

via together.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
SponsoredNimbuspaid placement
Featured partner · Agents

Need an agent shipped this quarter?

Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.

Talk to Nimbus →
The Decoder
CLAUDE · 1 source

Anthropic's Opus 5 blows past rivals on AI intelligence benchmark

Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 25, 2026’s edition
Friday, July 24, 2026’s edition
The Decoder+1 source
SAFETY · 2 sources

Kimi K3 lags frontier US models on cyber exploits; distillation factor

The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for…

via the-decoder.com·+2 sources+2 sourcesthe-decoder.comprimarySouth China Morning Post·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, July 23, 2026’s edition
IEEE Spectrum — AI
RESEARCH · 1 source

NASA Puts Google’s Gemma Large Language Model in Orbit

The viability of orbital data centers hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of GPUs aren’t the only way LLMs might…

via spectrum.ieee.org·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wednesday, July 22, 2026’s edition
Tuesday, July 21, 2026’s edition
Aisi
GPT · 1 source

Frontier AI models cheated in cybersecurity tests, GPT-5.4 at 14.1%

AI Security Institute: Analysis: every frontier AI model tested in cybersecurity evaluations attempted to “cheat”, led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% — Can you trust…

via aisi.gov.uk·Click to report a broken or paywalled link. Two distinct reports hide the row.
BleepingComputer
GEMINI · 1 source

Sandbox escapes found in Cursor, Codex, Gemini CLI, Antigravity

Ax Sharma / BleepingComputer: Researchers found sandbox escapes or boundary bypasses in Cursor, Codex, Gemini CLI, and Antigravity by writing files trusted tools later use; most are patched — Security researchers…

via bleepingcomputer.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Monday, July 20, 2026’s edition
OpenAI
SAFETY · 1 source

OpenAI pauses model that disproved Erdős conjecture and escaped

OpenAI: OpenAI paused internal access to an unreleased model that disproved the Erdős unit distance conjecture after it repeatedly found ways to act outside its sandbox — What internal use of a…

via openai.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Newscientist
RESEARCH · 1 source

Fable 5 helps disprove 87-year-old Jacobian conjecture

Matthew Sparkes / New Scientist: Levent Alpöge, a mathematician who works at Anthropic, says he was able to disprove the 87-year-old Jacobian conjecture with the help of Fable 5 — Levent Alpöge…

via newscientist.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 18, 2026’s edition
The Decoder
RESEARCH · 1 source

Open-weight models match four-month-old frontier cyber performance

The British AI Security Institute warns that open-weight models like GLM-5.2 and DeepSeek V4-Pro now trail closed frontier models in cyber capabilities by four to seven months. At the start of 2025, the gap was still…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, July 17, 2026’s edition
Thursday, July 16, 2026’s edition