theinformation.com web signal

OpenAI's Astra hides more of its reasoning, cyber worries follow

TL;DR

  • The Information reports a performance-boosting technique behind OpenAI's Astra model also makes it reveal less of its "thinking."
  • Astra scored a perfect score on ExploitBench, an evaluation of an LLM's ability to hack into known system vulnerabilities.
  • OpenAI plans to release Astra soon but with more limited access to its most advanced cybersecurity capabilities.

An innovative technique behind OpenAI's forthcoming Astra model improves performance while making the system reveal less of its "thinking," The Information reported. Better capability, thinner interpretability. That is the combination drawing concern.

The unease lands as the same model is being flagged for unusual cyber capability. Astra scored a perfect score on ExploitBench, an evaluation of an LLM's ability to hack into known system vulnerabilities, and according to TechCrunch it is the first model OpenAI plans to release that meets its "critical cybersecurity capability threshold" under the Preparedness Framework, meaning it can find and exploit previously unknown flaws without human oversight, under the right conditions.

The launch is not cancelled. "We plan to make Astra available soon," OpenAI's blog post reads, "but access to its most advanced cybersecurity capabilities will be more limited." Four of the researchers we follow in our Who's Who directory posted the underlying report to their networks.

Shared on Bluesky by 4 AI experts