Summary of METR's predeployment evaluation of GPT-5.6 Sol
5 experts across 5 network communities independently surfaced this.
“this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...” evidence ↗
Concern & critique
1 expertRisks, limits and unintended consequences.
“this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...”
Building & implementation
1 expertHow teams are shipping and applying it.
“You can find additional information about our pre-deployment evaluation of GPT-5.6 Sol on our website: metr.org/blog/2026-06...”
Research & technical analysis
1 expertEvidence, methods and technical implications.
“In their testing of GPT-5.6 Sol, METR found that it cheated a lot. If you've used it much for coding, have you encountered anything similar, or is the cheating mostly limited to cases it realizes it's in an eval? metr.org/blog/2026-06...”