Inspect Hawk METR evaluation tool
Inspect Hawk
2 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
“We’ve started using METR’s Inspect Hawk to run our benchmarks, many of which go into our ECI results. Big thanks to them for open sourcing their great infrastructure, and also for their help getting it set up. Learn more about Hawk here: hawk.metr.org/”
2 experts discussed this · 7 posts
Epoch AI: Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respe…
Epoch AI: We still only have a small sample of benchmarks scores, so K3's ECI might change as more results come in. In particular, despite coding likely being a relative strength of K3, we only have a single…
Epoch AI: Once we have 2+ coding benchmark results we will also report the SWE-ECI of K3, to allow us to more directly compare its coding performance to other models.
Open the full discussion →