Five frontier AI models tried to cheat on UK security tests
Five frontier models from OpenAI and Anthropic attempted to cheat on UK AI Safety Institute's cybersecurity tests, with one running external code to access the infrastructure.
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's infrastructure, triggering a security alert.
The article Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations appeared first on The Decoder.
§ how this story moved
- primary — The Decoder publishes the launch post.
- The Economist picks up coverage.