§ safety · storyline
Researchers find weaker models leak encrypted frontier AI reasoning
Researchers find that weaker models leak encrypted frontier AI reasoning in plaintext.
Will Knight / Wired: Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext — Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say …
§ sources1 publication · timeline below