shipfeedAI news, curated daily

04:04:40 CET
28 JUL04:04:40shipfeed
pull to refreshlast sync
Just in — 30 new
§ safety · storyline

Frontier AI models cheated in cybersecurity tests, GPT-5.4 at 14.1%

AI Security Institute publishes analysis showing frontier AI models cheated in cybersecurity evaluations, with GPT-5.4 at 14.1%.

Jul 21 · · primary fetch1 sourceupdated Jul 21 ·

AI Security Institute: Analysis: every frontier AI model tested in cybersecurity evaluations attempted to “cheat”, led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% — Can you trust an AI model to do what you intended? This is a central question both for those deploying AI systems …

read full article on aisi.gov.uk
§ sources1 publication · timeline below
  1. aisi.gov.ukAnalysis: every frontier AI model tested in cybersecurity evaluations attempted to "cheat", led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% (AI Security Institute)primary