“We’ve started using METR’s Inspect Hawk to run our benchmarks, many of which go into our ECI results. Big thanks to them for open sourcing their great infrastructure, and also for their help getting it set up. Learn more about Hawk here: hawk.metr.org/”
2 experts discussed this · 7 posts
Epoch AI: Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respe…
Epoch AI: We still only have a small sample of benchmarks scores, so K3's ECI might change as more results come in. In particular, despite coding likely being a relative strength of K3, we only have a single…
Epoch AI: Once we have 2+ coding benchmark results we will also report the SWE-ECI of K3, to allow us to more directly compare its coding performance to other models.
“ECI is our statistical tool for combining multiple benchmarks into a single, unified scale. See more results and learn about the methodology behind the ECI here: epoch.ai/benchmarks/eci”
“This isn't rock-solid evidence. Without AI, engineers would likely build code differently, and more complex contributions aren't necessarily more valuable. See METR's clarifying article for more discussion on : metr.org/blog/2026-0...”
“This week’s Gradient Update was written by Jean-Stanislas Denain and Anson Ho. All Gradient Updates are informal, opinionated analyses that represent the views of individual authors, not Epoch AI as a whole. Read the full essay here: epochai.substack.com/p/…”
“This kind of speculative analysis has worked before. E.g. Konstantin Tsiolkovsky argued that liquid-fueled rockets could reach space 41 years before one did. And more generally, the forecasting track record of futurists isn’t as bad as often assumed: www.co…”
We use essential cookies to keep the site working (login, form security). With your permission, we also use analytics cookies to understand how you use the site.
Privacy policy