We'll keep auditing and improving the ECI as it evolves. The updated ECI codebase remains open source at github.com/epoch-resea....
Epoch AI
Directory member with public evidence across AI business.
- AI signals
- 17 past 30d
- Sources
- 6 distinct domains
- Discussões
- 1 past 30d
- Latest signal
- 4d ago
Articles & links
Anthropic says Glasswing has surfaced 10k+ high- or critical-severity vulnerabilities so far (some remain unpublished). OpenAI's Daybreak program likely adds more. The spike in CVEs likely reflects this wave of AI-assisted discovery. Full Data Insight: epoch.ai/data-insigh...
For more details, check out our website! epoch.ai/MirrorCode
We’ve started using METR’s Inspect Hawk to run our benchmarks, many of which go into our ECI results. Big thanks to them for open sourcing their great infrastructure, and also for their help getting it set up. Learn more about Hawk here: hawk.metr.org/
GPT-5.6 Sol has been climbing Slay the Spire's Ascension ladder on our Twitch channel for a week, no human in the loop. This Thursday, Claude Opus 5 takes over the climb — live, with commentary. Thursday, July 30 · 12:30 PT twitch.tv/epochaiplays
Based on the July 2026 Epoch AI/Ipsos survey of 1,103 employed US adults, with 469 people having used AI for work in the past 7 days. Read the full data insight here: epoch.ai/data-insigh...
If you find these questions interesting and would like to help us answer them, apply to work at Epoch! Researchers: jobs.lever.co/epoch-ai/de... Engineers: jobs.lever.co/epoch-ai/d1... All careers: epoch.ai/about/careers
If you find these questions interesting and would like to help us answer them, apply to work at Epoch! Researchers: jobs.lever.co/epoch-ai/de... Engineers: jobs.lever.co/epoch-ai/d1... All careers: epoch.ai/about/careers
Gradient Updates are informal, opinionated analyses that represent the views of individual authors, not Epoch AI as a whole. Read the full essay here: epochai.substack.com/p/9-big-que...
This week’s Gradient Update was written by Campbell Hutcheson. All Gradient Updates are informal, opinionated analyses that represent the views of individual authors, not Epoch AI as a whole. Read the full essay here: epochai.substack.com/p/will-fina...
See our website for all the scores, as well as other benchmarks! epoch.ai/benchmarks/...
This Epoch AI/Ipsos survey was fielded July 10–19, 2026 (1,106 employed US adults) on KnowledgePanel, Ipsos’ probability-based panel. For the full analysis, read more here! epoch.ai/publication...
Recent commentary
AI appears to be finding software vulnerabilities at scale. In June 2026, 21 notable organizations disclosed ~1,500 high- and critical-severity CVEs, over 3.5× the previous monthly record set before Claude Mythos Preview's release.
How much more AI compute does a dollar buy each year? About 49%, or a doubling every 21 months, based on the chips actually bought each quarter from 2023 through 2025.
Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities.
Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna.
Introducing EBR-bench, our new benchmark to measure on-the-fly learning. AI repeatedly plays a challenging board game called Earthborne Rangers and tries to learn from its mistakes. So far: no signs of improvement.
Claude Fable 5 scores very well on FrontierMath: Tiers 1–4 (v2), reaching 87% on Tiers 1–3 and 88% on Tier 4. This continues a streak of Anthropic models improving rapidly at math.
We’ve launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far, and solving all of them would be an incredible mathematical feat. Thread with more.
How surprising should we find it that an internal OpenAI model was able to escape its restrictions and autonomously hack Hugging Face, all just to cheat on a cybersecurity benchmark? We have pulled together the public evidence on AI cyber capabilities in this thread:
DeepSeek-V4-Flash-0731 debuts with an ECI of 153, comparable to GLM 5.2 and roughly midway between Opus 4.5 and Opus 4.6. It's the second strongest open-weights model available today, behind only Kimi K3.
The end of the self-funded AI buildout? Hyperscaler cash capex is growing much faster than cash inflows. On current trends, they will be unable to fully fund the AI infrastructure buildout with cash from operations by the end of this year.
In Epoch AI's orbit
Center = Epoch AI. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.