shipfeedAI news, curated daily

04:03:07 CET
28 JUL04:03:07shipfeed
pull to refreshlast sync
Just in — 30 new
§ models · storyline

Kimi K3 lags frontier US models on cyber exploits; distillation factor

Moonshot AI's Kimi K3 scores 32 percent on cyber exploits versus 76 percent for leading U.S. models, with distillation potentially explaining the gap.

Jul 24 · · primary fetch2 sourcesupdated Jul 24 ·

The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks.

The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models. The article Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why appeared first on The Decoder.

read full article on the-decoder.com
§ sources2 publications · timeline below
  1. the-decoder.comKimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain whyprimary
  2. South China Morning PostChina’s Kimi K3 well behind US rivals in cyberattack ability, study shows

§ how this story moved

  1. primarySouth China Morning Post publishes the launch post.
  2. The Decoder picks up coverage.