Safety Agencies Rate AI 'Below Frontier' Same Week It Discovers 19 Zero-Days No One Knew About
LONDON — A joint evaluation released Thursday by the United Kingdom AI Safety Institute and its US counterpart found that China's Kimi K3 model "performs significantly below the most recent frontier cyber-capable models," scoring 32.2% on the agencies' ExploitBench assessment compared to a 76.2% average for top US frontier systems — a finding the agencies said triggers no additional safety reporting requirements under current guidelines.
The evaluation covered 41 known vulnerabilities. Kimi K3 achieved arbitrary code execution on zero of them.
Also this week, security researcher Chaofan Shou reported that a swarm of Kimi K3 agents had identified 19 previously unknown Redis vulnerabilities in approximately 90 minutes and produced a working remote code execution exploit in an additional 27. Redis shipped seven emergency security releases on July 23 across four major versions.
In a supplementary statement, the evaluation agencies said the results were not in conflict. ExploitBench, they explained, measures a model's ability to exploit documented vulnerabilities using partial scaffolding provided to the model. Discovery of novel zero-day vulnerabilities "falls outside the current evaluation framework," which was developed to reflect "real-world adversarial deployment scenarios as modeled in 2024."
An updated benchmark incorporating zero-day discovery capacity is in development and will be ready "no sooner than Q4 2027," the statement said. Until the new benchmark is published and validated, Kimi K3 "does not trigger enhanced oversight or additional reporting requirements."
The evaluation report does not reference the seven Redis emergency patches. The agencies confirmed the document had been finalized before the patches were issued and would be revisited in the next scheduled review cycle.
"The framework is working as intended," a spokesperson said.