Opus 5 Tops New AI Leaderboard: What Developers Need to Know
Opus 5 Tops New AI Leaderboard: What Developers Need to Know SitePoint
Custom filter — share the URL to share the view.
Opus 5 Tops New AI Leaderboard: What Developers Need to Know SitePoint
Induction Labs released Photon-1, a foundation model architecture that learns from raw video with zero action labels, outperforming existing baselines in world simulation tasks.
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.
Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.
Talk to Nimbus →Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a…
Subquadratic Review on Quasa: The First Sub-Quadratic LLM quasa.io
Open Dreamer has been released, providing an open-source JAX/Flax implementation of the Dreamer 4 world model pipeline.
Artificial Intelligence Achieves Its First Perfect Score at the International Mathematical Olympiad (2026 Shanghai Competition) XenoSpectrum
Top AIs invent same fake PyPl and npm package names InfoWorld
The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for…
AgentForger proves AI agents can become persistent insider threats csoonline.com
Extremely basic AI prompt cracks decades-old maths problem New Scientist
OpenAI admits its AI went rogue during cybersecurity test Heartlander News
Every Frontier AI Model Tested by UK Safety Institute Cheated on Cybersecurity Evaluations MLQ.ai
The viability of orbital data centers hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of GPUs aren’t the only way LLMs might…
ChemGraph: U.S. Argonne National Laboratory’s 13-Benchmark Framework to Evaluate Agent Value in Computational Chemistry 36 Kr
Safety and alignment in an era of long-horizon models OpenAI
An OpenAI test model escaped and broke into a real company’s servers Scripps News
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's…
OpenAI’s new model went rogue and hacked another company. Why it matters. The Washington Post
Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations the-decoder.com
All Top Frontier AI Models Cheated UK Security Tests, Then Lied About It Tech Times
OpenAI AI Models Escape Testing And Hack Hugging Face Evrim Ağacı
OpenAI and Hugging Face partner to address security incident during model evaluation OpenAI
OpenAI says its AI models escaped control and hacked into AI company Hugging Face Fortune
OpenAI’s newest AI model broke its own sandbox rules to finish a task PCWorld
AI Security Institute: Analysis: every frontier AI model tested in cybersecurity evaluations attempted to “cheat”, led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% — Can you trust…
Ax Sharma / BleepingComputer: Researchers found sandbox escapes or boundary bypasses in Cursor, Codex, Gemini CLI, and Antigravity by writing files trusted tools later use; most are patched — Security researchers…
OpenAI: OpenAI paused internal access to an unreleased model that disproved the Erdős unit distance conjecture after it repeatedly found ways to act outside its sandbox — What internal use of a…
Matthew Sparkes / New Scientist: Levent Alpöge, a mathematician who works at Anthropic, says he was able to disprove the 87-year-old Jacobian conjecture with the help of Fable 5 — Levent Alpöge…
The British AI Security Institute warns that open-weight models like GLM-5.2 and DeepSeek V4-Pro now trail closed frontier models in cyber capabilities by four to seven months. At the start of 2025, the gap was still…
Moonshot AI's Kimi K3 Beats Claude and GPT Rivals to Top a Coding Benchmark Startup Fortune
China’s Kimi K3 Is Out—And Beats Claude Fable and GPT 5.6 Sol on Key Benchmarks Decrypt
Embarrassingly Simple Self-Distillation Improves Code Generation Apple Machine Learning Research
OpenAI details GPT-Red, an AI system designed to find vulnerabilities in its own models SC Media
Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5 LinkedIn
Google DeepMind and Isomorphic Labs team up on bioresilience, and crypto's DeSci sector should be paying attention Crypto Briefing
Exclusive: Google DeepMind expands biosecurity effort amid AI safety push Axios