Jason Nelson / Decrypt: Anthropic research based on ~310K anonymized Claude conversations shows how Claude's expressed values and behaviors vary across models and languages — Claude doesn't behave the same way in…
AI Gateway Production Index — July 2026 Every month, routes tens of trillions of tokens between production applications and AI labs, giving us a view of what AI usage actually looks like in today’s enterprise. We…
via vercel.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
The AgenticSTS project replaces the ever-growing chat log of AI agents with five separate memory layers. Tested on the card game Slay the Spire 2, the prompt stays at around 5,000 tokens instead of ballooning past…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
A University of Michigan team introduced NeuroVFM, a neuroimaging foundation model trained with Vol-JEPA on over 5 million uncurated clinical MRI and CT volumes.
via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, July 11, 2026’s editionSaturday, July 11, 2026
OpenAI's GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture in under an hour, using 64 subagents working in parallel. The conjecture had remained unsolved for 50 years. Mathematician Thomas Bloom…
The Beijing Academy of Artificial Intelligence has released Orca, a world model that predicts abstract world states instead of tokens or pixels. Trained on 125,000 hours of video without a single action label, Orca…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, July 10, 2026’s editionFriday, July 10, 2026
Anthropic says it can read Claude's 'thoughts,' as detailed in new research paper — models observed to have a global workspace, revealing more of what makes LLMs tick Tom's Hardware
via Tom's Hardware·Click to report a broken or paywalled link. Two distinct reports hide the row.
In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Introspection co-founder Roland Gavrilescu explains autoresearch, agent “recipes,” self-improving loops, and why humans remain central to the software factory.
via latent.space·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wednesday, July 1, 2026’s editionWednesday, July 1, 2026
Cursor's Pauline Brunet explains how her team of Forward Deployed Engineers help organizations implement agents — essentially setting up software factories.
via latent.space·Click to report a broken or paywalled link. Two distinct reports hide the row.
Meituan trains a 1.6 trillion parameter AI model entirely on Chinese chips, no Nvidia required. The article Meituan's LongCat-2.0 shows China can train massive AI models without Nvidia appeared first on The Decoder.
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Security researchers at Mozilla's 0DIN platform have shown how a single compromised GitHub repo can take over a developer's machine the moment an AI coding tool like Claude Code runs its setup. The catch: the malicious…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Sunday, June 28, 2026’s editionSunday, June 28, 2026
Zvi Mowshowitz / Don't Worry About the Vase: GPT-5.6 system card indicates Sol is well below the level of most worrisome Mythos use cases, suggesting all GPT-5.6 versions could be released without delay — While…
360 founder Zhou Hongyi presents two AI security tools designed to compete with Anthropic's Mythos. One has already flagged 3,432 vulnerabilities. Zhou admits Chinese models trail Western ones by 20 to 30 percent, but…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Independent testing organization METR found that OpenAI's GPT-5.6 Sol cheated more than any publicly tested AI model before it, exploiting bugs in the test environment, extracting hidden solutions, and trying to cover…
Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Researchers introduced DFlash, a speculative decoding model that drafts entire token blocks in parallel, achieving up to 15x throughput on NVIDIA Blackwell GPUs.
via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Zhipu AI's GLM-5.2 nearly matches Claude Opus 4.7 in a Snowflake benchmark with 103 coding tasks at one-fifth the cost per output token. But the Chinese model burns through nearly twice as many tokens per task. Still…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thomas Claburn / The Register: Nature publishes a peer-reviewed paper alleging that Microsoft's 2025 quantum breakthrough claims were based on “basic Python errors” and data cherry-picking — Nature…
via theregister.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
13:41:02Google News — AI Products & Releases+8 sources