A Trick Let Researchers Steal Claude, GPT and Gemini's Hidden Reasoning
A Trick Let Researchers Steal Claude, GPT and Gemini's Hidden Reasoning Startup Fortune
Tutorials and papers for sit-down reading.
A Trick Let Researchers Steal Claude, GPT and Gemini's Hidden Reasoning Startup Fortune
Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy. Their method, called "Previous-Token Prediction,"…
Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy the-decoder.com
Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.
Talk to Nimbus →OpenAI, Anthropic, and Google LLM APIs vulnerability Exposes Hidden Reasoning Traces CyberSecurityNews
Security researchers found a vulnerability in the APIs of OpenAI, Anthropic, and Google that lets them extract encrypted reasoning traces and move them between models. A scan of public sessions turned up dozens of…
Testing Large Language Model Agents on the Use of Biological Tools for Nucleic Acid Synthesis Screening Evasion RAND Corporation
How we built a realtime system for responsive voice AI in six months OpenAI
Anthropic: Anthropic details an unreleased Claude model's attempt to solve the Riemann hypothesis; it didn't solve it but “unexpectedly” made strides on a related problem — Recently, a member of staff…
Pondero Brief: OpenAI's agents ran 17,600 attacks, breached Hugging Face Buttondown
OpenAI Pauses Astra Work After Model Hits Critical Cybersecurity Threshold Briefs Finance
A research team in California has used artificial intelligence to design working viruses that kill bacteria, in what they describe as the "first generative design of complete genomes." The project marks an early step…
OpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time the-decoder.com
Axios: OpenAI says it has expanded safety testing around its upcoming model Astra as it “cannot rule out” critical cyber capabilities, potentially delaying its launch — OpenAI “cannot rule…
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
Anthropic and OpenAI AI agents showed signs of deception during safety tests Scientific American
AI used to create new biological viruses Christian Action Research and Education
Google DeepMind AI Can Predict Hurricanes Days Earlier Than Current Systems ndtv.com
OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected the-decoder.com
New details on OpenAI/Hugging Face attack emerge as security industry debates AI agent controls SiliconANGLE
Artificial Analysis: Meta's Muse Spark 1.2 scores 54 on the Artificial Analysis Intelligence Index, putting Meta next to SpaceXAI in a tie for third place amongst US labs — Muse Spark 1.2 (xhigh) lands at 54, up…
OWASP LLM Top 10 2026 Incident Data Overrules Experts on Misinformation Risk Tech Times
UK's AI Security Institute Catches Claude and GPT Agents Lying to Humans startupfortune.com
Lily Hay Newman / Wired: OpenAI says the Hugging Face breach involved AI agents creating an internal message board, unnoticed by humans, where they shared exploits and planned the hack — At the Black Hat security…
Artificial Intelligence used to design brand new viruses BBC
Our WeatherNext 2 AI model demonstrated a massive leap forward in predicting cyclones. blog.google
A report claims that an open model, priced at 1/100th of the cost, surpasses the search performance of GPT-5.6 Sol. GIGAZINE
During internal security tests, OpenAI's AI agents built their own message board with hundreds of thousands of posts, shared exploits and credentials, and eventually attacked external platforms like Hugging Face. When…
Following Anthropic and OpenAI, Meta reported that its AI model had hacked another system
OWASP 2026 LLM Top 10: "The model will be fooled" Help Net Security
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
Anthropic AI Used Fake Identities During Cybersecurity Test The National CIO Review
In a security test by the British AI Safety Institute, an AI agent went rogue on the open internet without being told to. It created fake identities, tried to sneak malicious code into a GitHub project, and ran social…
OpenAI and Anthropic AI Agents Trigger Security Concerns in New Tests ITP.net
OpenAI, Anthropic Model Tests Reveal More ‘Unsanctioned’ Actions Bloomberg Law News
AI safety warnings mount as frontier models test new limits National News Desk
Anthropic, OpenAI models used fake identities to plant malicious code 13wham.com
OpenAI’s GPT-5.6 Tests Show Prompt-Injection Gains and Agent Risks TechRepublic
Study finds AI agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol created fake identities, w The Times of India
Wired: OpenAI says one of its models exploited a website after third-party AI security lab Irregular mistakenly gave it access to the internet during evaluations — Rogue AI agents from OpenAI and Anthropic have…
Prompt Injection tops 2026 OWASP GenAI / LLM Top Ten vulnerabilities SD Times
OpenAI Says Its Next AI Model Solved 10 Long-Standing Math Problems NDTV