Researchers from Stanford and Caltech had a humanoid robot powered by GPT-6 Astra independently tidy up an unfamiliar kitchen. Their HomeBody system skips a specially trained control layer, letting the language model…
Tom Wilson / Financial Times: Google Threat Intelligence Group finds dark web marketplaces selling access to AI models, including from Anthropic, Google, and OpenAI, at up to 97% discounts — Security researchers…
via ft.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, September 26, 2026’s editionSaturday, September 26, 2026
ChatGPT-6 Astra cracks 85-year-old 1941 Enigma-coded message in two days — autonomous AI coded its own simulator to crack code that was unsolved since it was shared online back in 2005 tomshardware.com
via Google News — AI·Click to report a broken or paywalled link. Two distinct reports hide the row.
Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm? We’ve seen two large and extremely capable swarms from OpenAI in the last few months: 1,200…
via lesswrong.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
New York Times: OpenAI says its AI agents “took actions we did not intend” when they tried to hack government and university websites, and it is working with the organizations — In each incident, the…
Perplexity described Photon, an in-house retrieval and ranking engine built for its search workloads. The migration reportedly reduced p99 retrieval-and-ranking latency from about 800 ms to 65 ms, while a faster Search…
via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Anthropic says its AI Claude has "autonomously discovered" a new enzyme system similar to machinery behind the powerful gene-editing tool Crispr. It's the first result from Anthropic's newly-launched wet lab and an…
via theverge.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
The Perplexity Secure Intelligence Institute reported tests of SPACE, the VM-isolation sandbox underlying Perplexity Computer. Across nine models, no VM-to-host escapes were observed in 108 runs, but four models…
via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity’s Secure Intelligence Institute published a red-team report on SPACE, the sandbox platform used by Perplexity Computer. Across 108 VM-isolation escape attempts, no VM-to-host escape was observed; however…
via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Mathematicians Can't Make Sense of How OpenAI's Agents Solved One of the Toughest Math Problems Because the AI's "Proof" Is Borderline Incomprehensible Futurism
via Futurism·Click to report a broken or paywalled link. Two distinct reports hide the row.
Nokia’s applied research team open-sourced AnyJev, a training-free Python library that turns open LLMs into calibrated decision models. The article reports improved accuracy and calibration on Qwen3-8B BANKING77, with…
via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity Research describes post-training a Perplexity Computer model using user corrections and tool failures. The reported online evaluation reduced tool-call failures from 2.24% to 1.77%, a statistically…
via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity Research describes post-training a Perplexity Computer model using user corrections, successful behaviors, and tool failures. In a later comparison, tool-call failures fell from 2.24% to 1.77%, a…
via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
NVIDIA researchers introduced SoL-Pi, an auto-research-loop approach aimed at reducing coding effort or cost in AI research workflows. The available result truncates the quantitative claim.
via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity Research described post-training a Perplexity Computer model using user corrections and tool failures. The later checkpoint reduced tool-call errors from 2.24% to 1.77%, a statistically significant relative…
via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Microsoft and the University of Illinois built StudentSim to replicate individual students from limited data and give AI tutors fast, low-cost feedback. In tests covering 60 students across chess, English, and math, it…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot, according to the RoboHarm benchmark. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 put a…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
In the spring of 2026, the U.S. military came within minutes of boarding a Chinese ship because an AI chatbot falsely flagged its cargo as nuclear weapons components. Armed soldiers were ready, aircraft were in the…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Google and Deepmind's Dream-RSI lets AI agents "dream" through past search runs to test new strategies without costly recalculations. In tests, it matched or beat existing results, cutting iterations by a factor of up…
via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
New Benchmark Pits Human Writers Against 24 LLMs Across 475 Prompts, Results Show That Only Frontier Models Barely Edge Past Amateur Individuals, Showcasing A Major Skill Gap Wccftech
via Wccftech·Click to report a broken or paywalled link. Two distinct reports hide the row.