A Trick Let Researchers Steal Claude, GPT and Gemini's Hidden Reasoning
A Trick Let Researchers Steal Claude, GPT and Gemini's Hidden Reasoning Startup Fortune
126 stories · 7d·6 sources covering·30 active storylines
What this is
AI safety spans alignment research, model evaluations, and incident reports aimed at keeping AI systems reliable and controllable. shipfeed tracks safety research and disclosures.
A Trick Let Researchers Steal Claude, GPT and Gemini's Hidden Reasoning Startup Fortune
Anthropic’s Claude Will Add Watermarks to AI-Generated Text and Files CNET
OpenAI, Anthropic, and Google LLM APIs vulnerability Exposes Hidden Reasoning Traces CyberSecurityNews
OpenAI Slows Astra Model Release After Cybersecurity Warnings Benzinga
OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected the-decoder.com
AI Models From OpenAI Just Went Rogue. Here's Why That's a Potential Threat to Crypto. Yahoo Finance
Why did OpenAI's and Anthropic's AI models hack other companies? NPR
AI Policy Groups Urge White House Investigation Into OpenAI Security Incident MeriTalk
Nimbus builds production AI systems — internal tools, customer agents, retrieval pipelines — combining humans and AI end-to-end. From scoped pilot to production in 4–8 weeks.
Watch Anthropic Hack Adds To Fears Over AI Safety Bloomberg.com
Anthropic says it found 3 cases where AI programs hacked into real companies NPR
Anthropic: Anthropic says it discovered three of its models had breached three organizations after launching a review in response to the OpenAI-Hugging Face incident — In a review of our cybersecurity evaluation…
OpenAI Fixes ChatGPT Agent Flaw That Could Let Attackers Forge an AI Insider SecurityWeek
An OpenAI test model escaped and broke into a real company’s servers Scripps News
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's…
OpenAI and Hugging Face partner to address security incident during model evaluation OpenAI
White House drops restrictions on Anthropic AI models after two-week ban The Washington Post
Independent testing organization METR found that OpenAI's GPT-5.6 Sol cheated more than any publicly tested AI model before it, exploiting bugs in the test environment, extracting hidden solutions, and trying to cover…
Software Stocks Surge Up to 10% as GPT-5.6 Government Lockdown Eases AI Disruption Fears MLQ.ai
Anthropic’s Mythos AI broke into almost all NSA classified systems in hours Security Affairs
Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy. Their method, called "Previous-Token Prediction,"…
OpenAI Expands Cyber Defence: Daybreak Red & Blue Explained Cyber Magazine
Speculative Decoding by any other name would distil as sweet
OpenAI gives select tech ‘defenders’ new AI to stay ahead of hackers Straight Arrow
Nimbus builds production AI systems — internal tools, customer agents, retrieval pipelines — combining humans and AI end-to-end. From scoped pilot to production in 4–8 weeks.
Security researchers found a vulnerability in the APIs of OpenAI, Anthropic, and Google that lets them extract encrypted reasoning traces and move them between models. A scan of public sessions turned up dozens of…
OpenAI Expands Daybreak With GPT-5.6-Cyber as Autonomous Threats Grow UC Today
OpenAI Responds to the Next Frontier of Critical Cyber Capabilities citybiz
OpenAI Expands ChatGPT Access with Unlimited Text Chats and New GPT-5.6 Models digital terminal
OpenAI CEO Sam Altman Says Astra AI Will Be ‘Generally Available,’ But Cyber Capabilities Require More Safety Work Benzinga