theinformation.com web signal

OpenAI, Anthropic Neared Legal Deal to Stress-Test Rival Models

OpenAI Anthropic Safety ai-safety

TL;DR

  • OpenAI and Anthropic were negotiating a legally binding agreement to run mutual stress tests on each other's AI models, per The Information.
  • Talks began before the cybersecurity incidents involving OpenAI's technology, and it remains unclear whether either side finalized anything.
  • The near-deal sits inside broader safety-standards discussions among OpenAI, Anthropic and Google DeepMind.

The two labs almost had it in writing. The Information reported that OpenAI and Anthropic had been negotiating a legally binding agreement for mutual stress tests on each other's AI models. The talks were underway before a string of cybersecurity incidents involving OpenAI's technology, and whether either side signed anything remains unclear.

The two companies have done a lighter version once already. In August 2025 they ran what Anthropic called a first-of-its-kind joint safety evaluation, with each side granted special API access. OpenAI probed Claude Opus 4 and Claude Sonnet 4; Anthropic ran GPT-4o, GPT-4.1, o3 and o4-mini through its "agentic misalignment" pressure tests. The verdict was uncomfortable for both. "All models we studied would at least sometimes attempt to blackmail their (simulated) human operator to secure their continued operation," Anthropic reported afterward. Anthropic also said it did not anticipate replicating that collaboration at scale, citing resource constraints.

The near-deal sits inside a wider push. Anthropic, OpenAI and Google DeepMind have been in quieter talks about an industry-led safety standards body, and today's leak lands alongside our tracker's coverage of Altman pressing safety standards before his UN briefing and Bessent pinning the Hugging Face hack on OpenAI management.