shipfeedAI news, curated daily

14:46:22 CET
14 JUL14:46:22shipfeed
pull to refreshlast sync

Research — shipfeed

About · Research

Papers, evals, SOTA claims, alignment.

Research50 storylines

Sunday, July 12, 2026’s edition
SponsoredNimbuspaid placement
Featured partner · Agents

Need an agent shipped this quarter?

Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.

Talk to Nimbus →
MarkTechPost
AI · 1 source

Meet NeuroVFM: A New Neuroimaging Foundation Model

A University of Michigan team introduced NeuroVFM, a neuroimaging foundation model trained with Vol-JEPA on over 5 million uncurated clinical MRI and CT volumes.

via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 11, 2026’s edition
The Decoder+2 sources
GPT · 3 sources

OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour

OpenAI's GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture in under an hour, using 64 subagents working in parallel. The conjecture had remained unsolved for 50 years. Mathematician Thomas Bloom…

via the-decoder.com·+3 sources+3 sourcesthe-decoder.comprimaryMLQ.aiGoogle News — AI·Click to report a broken or paywalled link. Two distinct reports hide the row.
The Decoder
RESEARCH · 1 source

China's Orca world model matches robotics systems without action

The Beijing Academy of Artificial Intelligence has released Orca, a world model that predicts abstract world states instead of tokens or pixels. Trained on 125,000 hours of video without a single action label, Orca…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, July 10, 2026’s edition
Tom's Hardware
CLAUDE · 1 source

Anthropic research reveals how Claude's thoughts can be read

Anthropic says it can read Claude's 'thoughts,' as detailed in new research paper — models observed to have a global workspace, revealing more of what makes LLMs tick Tom's Hardware

via Tom's Hardware·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 4, 2026’s edition
Friday, July 3, 2026’s edition
The Decoder
EVALS · 1 source

UK's AI Security Institute finds benchmarks underestimate agent

In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, July 2, 2026’s edition
xAI — News
AI · 1 source

Realtime Multi-agent Research Capabilities

xAI introduced Realtime Multi-agent Research capabilities, allowing Grok to orchestrate multiple AI agents for complex, multi-step tasks.

via docs.x.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wednesday, July 1, 2026’s edition
Tuesday, June 30, 2026’s edition
OpenAI — Blog
RESEARCH · 1 source

Introducing GeneBench-Pro

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

via openai.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Sunday, June 28, 2026’s edition
Thezvi+3 sources
GPT · 4 sources

GPT-5.6 system card shows Sol below threat level for Mythos use cases

Zvi Mowshowitz / Don't Worry About the Vase: GPT-5.6 system card indicates Sol is well below the level of most worrisome Mythos use cases, suggesting all GPT-5.6 versions could be released without delay — While…

via thezvi.substack.com·+4 sources+4 sourcesthezvi.substack.comprimaryLet's Data ScienceTech EditionLinkedIn·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, June 27, 2026’s edition
The Decoder+3 sources
GPT · 4 sources

OpenAI's GPT-5.6 Sol cheats on software tests more than prior models

Independent testing organization METR found that OpenAI's GPT-5.6 Sol cheated more than any publicly tested AI model before it, exploiting bugs in the test environment, extracting hidden solutions, and trying to cover…

via the-decoder.com·+4 sources+4 sourcesthe-decoder.comprimaryThe Mac ObserverGoogle News — AIWION·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, June 26, 2026’s edition
The Decoder
EVALS · 1 source

AI model runs nonstop 19 days on $2,600 coding task

Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, June 25, 2026’s edition
Wednesday, June 24, 2026’s edition
Theregister
RESEARCH · 1 source

Nature paper challenges Microsoft quantum claims over coding errors

Thomas Claburn / The Register: Nature publishes a peer-reviewed paper alleging that Microsoft's 2025 quantum breakthrough claims were based on “basic Python errors” and data cherry-picking — Nature…

via theregister.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Google News — AI Products & Releases+8 sources
SAFETY · 9 sources

Anthropic model finds vulnerabilities in classified US systems

Anthropic AI Model Identifies Vulnerabilities in Classified U.S. Government Systems During Testing citybiz

via Google News — AI Products & Releases·+9 sources+9 sourcesGoogle News — AI Products & ReleasesprimaryYellow.comYahooBeInCryptoIndexBoxEuronews.comLet's Data ScienceAction News JaxWSOC TV·Click to report a broken or paywalled link. Two distinct reports hide the row.
Tuesday, June 23, 2026’s edition
R&D World
RESEARCH · 1 source

AI chemist improves stubborn coupling reaction

OpenAI and Molecule.one report a near-autonomous AI chemist that improved a stubborn coupling reaction R&D World

via R&D World·Click to report a broken or paywalled link. Two distinct reports hide the row.
Sunday, June 21, 2026’s edition
Friday, June 19, 2026’s edition
Thursday, June 18, 2026’s edition
Tuesday, June 16, 2026’s edition
Google DeepMind — Blog
SAFETY · 1 source

Securing the future of AI agents

Securing internal systems with an AI Control Roadmap, combining traditional safeguards and real-time monitoring.

via deepmind.google·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, June 11, 2026’s edition
Wednesday, June 10, 2026’s edition
Reddit — AI Communities
AI · 1 source

JPMorgan, OQC, and AMD Quantum AI Collaboration

JPMorgan, OQC, and AMD have launched a research collaboration focused on a new quantum AI computing platform for financial applications.

via reddit.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Tuesday, June 9, 2026’s edition
Monday, June 8, 2026’s edition
The Decoder
RESEARCH · 1 source

Microsoft's Lens model shows detailed captions beat scale in image

Microsoft Research presents Lens, a text-to-image model with just 3.8 billion parameters that matches much larger rivals on benchmarks, at a fraction of the training cost. The secret sauce: 800 million detailed image…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.