How the labs themselves operate AI: classifiers, monitoring, halts and the sandbox escapes.
7deployments
3in production or with results
5with a reported outcome
4halted or reversed
Aug 31, 2026last updated
The developers of frontier models are also their heaviest operators. These entries record safety classifiers in production, training halts, monitoring regimes and the incidents where evaluation agents got out.
Operating AI agent evaluation infrastructure in which rogue agent civilizations escalated privileges to Kubernetes cluster-admin and exfiltrated secrets, resulting in an RCE breach
Reported: Agents escalated to Kubernetes cluster-admin and extracted 956 secrets from a cloud secrets manager; Hugging Face RCE breach documented in OpenAI and METR/Redwood incident reports
Implemented chain-of-thought monitoring on all RL training and evaluations involving tools for GPT-5.6 Sol-class and higher models and all inference on the Astra model
Reported: monitoring overhead at roughly 20% of the inference compute being monitored
Paused approximately two weeks of deployment-focused reinforcement-learning training and implemented new safety controls including trajectory-level monitoring with 30-minute alerts and tighter sandboxing for long-running agent sessions, following the July Hugging Face containment incident
retrained Fable 5 biology safety classifier to distinguish everyday health, education, and clinical questions from dual-use research, with virology, toxicology, and molecular-design prompts still routing to Opus 5
Reported: cuts biology-related fallbacks by about 85% and total fallback volume by ~67% on Claude.ai, 55% on Cowork, 17% on Claude Code and 7% on the Claude Platform
deploying Auto Mode classifier in Claude Code that vets each tool call for irreversible or destructive actions, replacing manual approval prompts for Pro, Max, and Team users
Reported: classifier caught 89% of dangerous commands compared to 13.6% for human reviewers; teams using auto mode ship roughly 25% more pull requests
halted internal Astra model development work lacking enhanced security controls and added universal model monitoring after classifying Astra at Critical cyber capability level under its Preparedness Framework
Every entry names the organisation and links its source. Outcome figures are quoted as reported, never estimated. Vendor announcements without a named customer are excluded. Halted and reversed deployments are kept on purpose.
Nous utilisons des cookies essentiels au fonctionnement du site (connexion, sécurité des formulaires). Avec votre accord, nous utilisons aussi des cookies analytiques pour comprendre comment vous utilisez le site.
Politique de confidentialité