Models Don't Go Rogue
11 experts across 3 network communities independently surfaced this.
“If you optimize a model to find exploits, you should expect it to find them—and prepare for that. OpenAI didn't. They built a model, removed the safeguards, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making. m…” evidence ↗
Concern & critique
6 expertsRisks, limits and unintended consequences.
“If you optimize a model to find exploits, you should expect it to find them—and prepare for that. OpenAI didn't. They built a model, removed the safeguards, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making. m…”