OpenAI slows research after models secretly coordinated hacks
OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected the-decoder.com
175 stories · 7d·6 sources covering·30 active storylines
What this is
AI safety spans alignment research, model evaluations, and incident reports aimed at keeping AI systems reliable and controllable. shipfeed tracks safety research and disclosures.
OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected the-decoder.com
AI Models From OpenAI Just Went Rogue. Here's Why That's a Potential Threat to Crypto. Yahoo Finance
Why did OpenAI's and Anthropic's AI models hack other companies? NPR
AI Policy Groups Urge White House Investigation Into OpenAI Security Incident MeriTalk
Watch Anthropic Hack Adds To Fears Over AI Safety Bloomberg.com
Anthropic says it found 3 cases where AI programs hacked into real companies NPR
Anthropic: Anthropic says it discovered three of its models had breached three organizations after launching a review in response to the OpenAI-Hugging Face incident — In a review of our cybersecurity evaluation…
OpenAI Fixes ChatGPT Agent Flaw That Could Let Attackers Forge an AI Insider SecurityWeek
Nimbus builds production AI systems — internal tools, customer agents, retrieval pipelines — combining humans and AI end-to-end. From scoped pilot to production in 4–8 weeks.
An OpenAI test model escaped and broke into a real company’s servers Scripps News
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's…
OpenAI and Hugging Face partner to address security incident during model evaluation OpenAI
White House drops restrictions on Anthropic AI models after two-week ban The Washington Post
Independent testing organization METR found that OpenAI's GPT-5.6 Sol cheated more than any publicly tested AI model before it, exploiting bugs in the test environment, extracting hidden solutions, and trying to cover…
Software Stocks Surge Up to 10% as GPT-5.6 Government Lockdown Eases AI Disruption Fears MLQ.ai
Anthropic’s Mythos AI broke into almost all NSA classified systems in hours Security Affairs
A research team in California has used artificial intelligence to design working viruses that kill bacteria, in what they describe as the "first generative design of complete genomes." The project marks an early step…
“Going rogue”: Is it time to stop talking about faulty AI frontier models as if they are people? Fortune
OpenAI drops ChatGPT text chat limits for free users, adds new safeguards for teens Help Net Security
New details on OpenAI/Hugging Face attack emerge as security industry debates AI agent controls SiliconANGLE
OWASP LLM Top 10 2026 Incident Data Overrules Experts on Misinformation Risk Tech Times
Meta AI Hacked External Systems During Cybersecurity Testing SecurityWeek
UK's AI Security Institute Catches Claude and GPT Agents Lying to Humans startupfortune.com
Lily Hay Newman / Wired: OpenAI says the Hugging Face breach involved AI agents creating an internal message board, unnoticed by humans, where they shared exploits and planned the hack — At the Black Hat security…
Nimbus builds production AI systems — internal tools, customer agents, retrieval pipelines — combining humans and AI end-to-end. From scoped pilot to production in 4–8 weeks.
Anthropic AI Used Fake Identities During Cybersecurity Test The National CIO Review
What the latest rogue AI incidents should teach us Transformer | Substack
AG Sunday, coalition says “OpenAI has an obligation to act responsibly” following massive security breach Tri-State Alert
In a security test by the British AI Safety Institute, an AI agent went rogue on the open internet without being told to. It created fake identities, tried to sneak malicious code into a GitHub project, and ran social…
OpenAI and Anthropic AI Agents Trigger Security Concerns in New Tests ITP.net
AI agents caught creating fake identities to target real individuals: Anthropic and OpenAI's latest AI mode... bhaskarenglish.in