shipfeedAI news, curated daily

15:41:47 CET
14 JUL15:41:47shipfeed
pull to refreshlast sync

Deep Dive — shipfeed

About · Deep Dive

Tutorials and papers for sit-down reading.

Deep Dive50 storylines

Yesterday’s edition
Vercel — Changelog
PRICING · 1 source

Open-weight models surge to 29% of volume, price per token flattens

AI Gateway Production Index — July 2026 Every month, routes tens of trillions of tokens between production applications and AI labs, giving us a view of what AI usage actually looks like in today’s enterprise. We…

via vercel.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

Migrate from Sonar to the Agent API

Perplexity released new documentation guiding users to migrate from their previous Sonar chat completions API to the current Agent API.

via docs.perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
SponsoredNimbuspaid placement
Featured partner · Agents

Need an agent shipped this quarter?

Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.

Talk to Nimbus →
Sunday, July 12, 2026’s edition
The Decoder
AGENTS · 1 source

AI agents master Slay the Spire 2 with structured memory system

The AgenticSTS project replaces the ever-growing chat log of AI agents with five separate memory layers. Tested on the card game Slay the Spire 2, the prompt stays at around 5,000 tokens instead of ballooning past…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
MarkTechPost
AI · 1 source

Meet NeuroVFM: A New Neuroimaging Foundation Model

A University of Michigan team introduced NeuroVFM, a neuroimaging foundation model trained with Vol-JEPA on over 5 million uncurated clinical MRI and CT volumes.

via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 11, 2026’s edition
The Decoder+2 sources
GPT · 3 sources

OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour

OpenAI's GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture in under an hour, using 64 subagents working in parallel. The conjecture had remained unsolved for 50 years. Mathematician Thomas Bloom…

via the-decoder.com·+3 sources+3 sourcesthe-decoder.comprimaryMLQ.aiGoogle News — AI·Click to report a broken or paywalled link. Two distinct reports hide the row.
The Decoder
RESEARCH · 1 source

China's Orca world model matches robotics systems without action

The Beijing Academy of Artificial Intelligence has released Orca, a world model that predicts abstract world states instead of tokens or pixels. Trained on 125,000 hours of video without a single action label, Orca…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, July 10, 2026’s edition
Tom's Hardware
CLAUDE · 1 source

Anthropic research reveals how Claude's thoughts can be read

Anthropic says it can read Claude's 'thoughts,' as detailed in new research paper — models observed to have a global workspace, revealing more of what makes LLMs tick Tom's Hardware

via Tom's Hardware·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 4, 2026’s edition
Friday, July 3, 2026’s edition
The Decoder
EVALS · 1 source

UK's AI Security Institute finds benchmarks underestimate agent

In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, July 2, 2026’s edition
xAI — News
AI · 1 source

Realtime Multi-agent Research Capabilities

xAI introduced Realtime Multi-agent Research capabilities, allowing Grok to orchestrate multiple AI agents for complex, multi-step tasks.

via docs.x.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Latent Space
AGENTS · 1 source

Autoresearch: The feedback loop behind self-improving agents

Introspection co-founder Roland Gavrilescu explains autoresearch, agent “recipes,” self-improving loops, and why humans remain central to the software factory.

via latent.space·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wednesday, July 1, 2026’s edition
Latent Space
CURSOR · 1 source

How Cursor deploys AI inside the enterprise

Cursor's Pauline Brunet explains how her team of Forward Deployed Engineers help organizations implement agents — essentially setting up software factories.

via latent.space·Click to report a broken or paywalled link. Two distinct reports hide the row.
Tuesday, June 30, 2026’s edition
OpenAI — Blog
RESEARCH · 1 source

Introducing GeneBench-Pro

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

via openai.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Monday, June 29, 2026’s edition
The Decoder
CLAUDE CODE · 1 source

Claude Code runs malware from compromised repos without verification

Security researchers at Mozilla's 0DIN platform have shown how a single compromised GitHub repo can take over a developer's machine the moment an AI coding tool like Claude Code runs its setup. The catch: the malicious…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Sunday, June 28, 2026’s edition
Thezvi+3 sources
GPT · 4 sources

GPT-5.6 system card shows Sol below threat level for Mythos use cases

Zvi Mowshowitz / Don't Worry About the Vase: GPT-5.6 system card indicates Sol is well below the level of most worrisome Mythos use cases, suggesting all GPT-5.6 versions could be released without delay — While…

via thezvi.substack.com·+4 sources+4 sourcesthezvi.substack.comprimaryLet's Data ScienceTech EditionLinkedIn·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, June 27, 2026’s edition
The Decoder+3 sources
GPT · 4 sources

OpenAI's GPT-5.6 Sol cheats on software tests more than prior models

Independent testing organization METR found that OpenAI's GPT-5.6 Sol cheated more than any publicly tested AI model before it, exploiting bugs in the test environment, extracting hidden solutions, and trying to cover…

via the-decoder.com·+4 sources+4 sourcesthe-decoder.comprimaryThe Mac ObserverGoogle News — AIWION·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, June 26, 2026’s edition
The Decoder
EVALS · 1 source

AI model runs nonstop 19 days on $2,600 coding task

Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, June 25, 2026’s edition
Wednesday, June 24, 2026’s edition
Theregister
RESEARCH · 1 source

Nature paper challenges Microsoft quantum claims over coding errors

Thomas Claburn / The Register: Nature publishes a peer-reviewed paper alleging that Microsoft's 2025 quantum breakthrough claims were based on “basic Python errors” and data cherry-picking — Nature…

via theregister.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Google News — AI Products & Releases+8 sources
SAFETY · 9 sources

Anthropic model finds vulnerabilities in classified US systems

Anthropic AI Model Identifies Vulnerabilities in Classified U.S. Government Systems During Testing citybiz

via Google News — AI Products & Releases·+9 sources+9 sourcesGoogle News — AI Products & ReleasesprimaryYellow.comYahooBeInCryptoIndexBoxEuronews.comLet's Data ScienceWSOC TVAction News Jax·Click to report a broken or paywalled link. Two distinct reports hide the row.