METR

Why they matter

Directory member with public evidence across Evaluation & benchmarks.

AI signals
1
past 30d
Sources
1
distinct domains
Discusiones
0
past 30d
Latest signal
7d ago
View every signal from METR →
METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.

Articles & links

Recent commentary

OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon.

View on Bluesky · ♥ 18 ↻ 0 ↩ 1 · 32d ago

Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.

View on Bluesky · ♥ 7 ↻ 2 ↩ 1 · 7d ago

In METR's orbit

Center = METR. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.