You can find additional information about our pre-deployment evaluation of GPT-5.6 Sol on our website: https://t.co/jvGwK5ljp7
METR
Directory member with public evidence across Evaluation & benchmarks.
- AI signals
- 1 past 30d
- Sources
- 1 distinct domains
- Discussions
- 0 past 30d
- Latest signal
- 7d ago
Articles & links
See the post for more, including: (1) a sketch of what we know about AI-assisted R&D; (2) alternative metrics for optimization ability; (3) estimated returns to human labor in NanoGPT; (4) details on NanoGPT agent runs. metr.org/blog/2026-07...
Recent commentary
OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon.
Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.
In METR's orbit
Center = METR. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.