Moonshot's Kimi K3 lags US frontier models on cyber exploits
TL;DR
- Kimi K3 scored 32% on ExploitBench versus 76% for US frontier models, failing to achieve arbitrary code execution on all 41 tasks.
- Kimi K3 became the first non-US model to solve the TLO simulated corporate network-attack benchmark, completing it once in 10 attempts.
- Kimi K3's safety safeguards did not block it from attempting cyber exploit development or offensive operations, with no jailbreak required.
The interesting number in this joint UK-US evaluation is not the headline 32.2%, it is the zero. Moonshot's newly released Kimi K3, as reported by the South China Morning Post, failed to achieve arbitrary code execution on any of the 41 Chrome V8 vulnerabilities in the ExploitBench suite. Leading unnamed US frontier models succeeded on 20 of 41 on average. On the same benchmark Kimi K3 scored 32.2%, ahead of Zhipu's GLM-5.2 at 24.4% but well behind the 76.2% US average.
The assessment, published by the UK AI Security Institute and the US Center for AI Standards and Innovation, also ran a 32-step simulated attack against a corporate network called 'The Last Ones'. Kimi K3 reached step 17 on average; the most cyber-capable US models reached 28.5. Full-path success came on one attempt in ten, within a 100M token limit.
The safety finding is the other one worth chewing on. The evaluators write that Kimi K3's safeguards 'did not prevent it from attempting cyber exploit development or offensive cyber operations' during their tests. So the model is meaningfully less capable than the American frontier on this workload, and it will not refuse the request. That combination, weaker on capability but essentially free on refusal, is what makes the open-weight distribution question harder rather than easier. Moonshot's open-weight release is slated for July 27.
One reading you will see elsewhere is that Kimi K3 leaned heavily on distilled outputs from US models, and that Anthropic-style safety classifiers stripped out the offensive-cyber examples in the process, leaving a strong generalist with a specific hole. The Decoder makes that argument. Treat it as a hypothesis, not a finding in the AISI write-up.
Bear in mind this is a preliminary evaluation on two tasks, the 'top US models' set is unnamed, and a single release is a snapshot. Neither the AISI note nor the SCMP piece says whether Moonshot will ship stronger refusals before the open-weight drop, or how quickly the capability gap narrows if the distillation story turns out to be right. For defenders the near-term read is calmer than the headline number suggested. For policymakers weighing open-weight controls, the guardrail gap is the harder problem.
What others are reporting
-
UK AISI Read →
First-party official report. Provides methodology (Item Response Theory), benchmark-level breakdown, and the verbatim safety conclusion on K3's default-assist posture.
Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during UK AISI / CAISI's evaluations.
-
NIST Read →
US government amplification. Notes TLO benchmark solves are no longer exclusive to US models, framing K3 as a marker of widening international cyber capability.
Solves of TLO are no longer exclusive to a small set of models.
-
The Decoder Read →
Offers a technical explanation for the capability gap: distillation from safety-aligned Western models would have stripped offensive cyber techniques filtered at the source.
Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so.
Originally reported by scmp.com
Read the original article →Original headline: UK AISI, US CAISI Joint Study Finds Kimi K3 Trails US Frontier Models on Cyber Exploits, 32% vs 76%