scmp.com web signal

Moonshot's Kimi K3 lags US frontier models on cyber exploits

TL;DR

  • A joint UK AISI and US CAISI evaluation scored Kimi K3 at 32.2% on ExploitBench, versus a 76.2% average for top US frontier models.
  • Kimi K3 failed to achieve arbitrary code execution on any of 41 Chrome V8 vulnerabilities; leading US models succeeded on 20 of 41 on average.
  • The evaluators found Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during testing.

The interesting number in this joint UK-US evaluation is not the headline 32.2%, it is the zero. Moonshot's newly released Kimi K3, as reported by the South China Morning Post, failed to achieve arbitrary code execution on any of the 41 Chrome V8 vulnerabilities in the ExploitBench suite. Leading unnamed US frontier models succeeded on 20 of 41 on average. On the same benchmark Kimi K3 scored 32.2%, ahead of Zhipu's GLM-5.2 at 24.4% but well behind the 76.2% US average.

The assessment, published by the UK AI Security Institute and the US Center for AI Standards and Innovation, also ran a 32-step simulated attack against a corporate network called 'The Last Ones'. Kimi K3 reached step 17 on average; the most cyber-capable US models reached 28.5. Full-path success came on one attempt in ten, within a 100M token limit.

The safety finding is the other one worth chewing on. The evaluators write that Kimi K3's safeguards 'did not prevent it from attempting cyber exploit development or offensive cyber operations' during their tests. So the model is meaningfully less capable than the American frontier on this workload, and it will not refuse the request. That combination, weaker on capability but essentially free on refusal, is what makes the open-weight distribution question harder rather than easier. Moonshot's open-weight release is slated for July 27.

One reading you will see elsewhere is that Kimi K3 leaned heavily on distilled outputs from US models, and that Anthropic-style safety classifiers stripped out the offensive-cyber examples in the process, leaving a strong generalist with a specific hole. The Decoder makes that argument. Treat it as a hypothesis, not a finding in the AISI write-up.

The honest caveat is this is a preliminary evaluation on two tasks, the 'top US models' set is unnamed, and a single release is a snapshot. What the reporting does not tell you is whether Moonshot will ship stronger refusals before the open-weight drop, or how quickly the capability gap narrows if the distillation story turns out to be right. For defenders the near-term read is calmer than the headline number suggested. For policymakers weighing open-weight controls, the guardrail gap is the harder problem.