⚡ 455 h early
OpenAI limits GPT-5.6 rollout government request
Summary of METR's predeployment evaluation of GPT-5.6 Sol
5 experts across 5 network communities independently surfaced this.
5 experts
5 communities
1 sources clustered
“Some quotes about Sol cheating metr.org/blog/2026-06... 👇”
“this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...”
2 experts discussed this · 6 posts
Tim Kellogg: this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...
Ted Underwood: this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...
Tim Kellogg: lol right? so badly wanted a hacker that they got one, and it’s not the kind of employee you want around
Open the full discussion →