Summary of METR's predeployment evaluation of GPT-5.6 Sol
5 experts across 5 network communities independently surfaced this.
“Some quotes about Sol cheating metr.org/blog/2026-06... 👇”
“this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...”