openai.com web signal

OpenAI pauses frontier RL training over Astra cyber threshold

TL;DR

  • OpenAI paused RL training on its latest deployment models for two weeks after flagging Astra as possibly meeting the Critical cyber threshold.
  • Monitoring safeguards now consume around 20% of inference compute, targeting an alert within 30 minutes of concerning activity.
  • All Astra workloads require the strictest safeguards; the largest planned frontier RL run is on hold pending smaller-scale evaluations.

OpenAI has paused reinforcement learning training on its latest deployment models for two weeks after preliminary evidence that its upcoming model Astra "may meet the Critical cybersecurity capability threshold" under the company's Preparedness Framework. In its post, OpenAI says it "temporarily slowed" the pace of scaling to harden research environments, expand monitoring coverage, and rerun evaluations.

Critical is the top tier of OpenAI's rating system. All Astra workloads now require the strictest level of security safeguards, and the largest planned frontier RL run is on hold pending smaller-scale training and evaluations.

The monitoring layer aims to issue an alert within 30 minutes of concerning activity, at an estimated overhead of 20% of inference compute, a cost the company notes "varies substantially across training and evaluation workloads." OpenAI frames its safety approach around three lines: monitoring, alignment, and security. "We now require stronger evidence of aligned behavior throughout all of training," the post reads.

The Decoder reports OpenAI attributed the slowdown to a Hugging Face security incident and "rapid progress in our internal research." Four experts in our Who's Who directory shared the piece the same week.

Shared on Bluesky by 4 AI experts