wsj.com web signal

Fired OpenAI safety trio urges board to halt opaque-reasoning AI

OpenAI Safety Sam Altman ai-business

TL;DR

  • Three former OpenAI safety researchers wrote to the board urging the company and its rivals to stop building models whose reasoning humans cannot audit.
  • Fired on October 1 for allegedly mishandling sensitive infrastructure data, the three include lead authors of the 'Chain of Thought Monitorability' paper.
  • The letter claims the firings have created a 'chilling' atmosphere for remaining safety staff; OpenAI denies the terminations relate to safety concerns.

Three safety researchers OpenAI dismissed on October 1 have written to the company's board asking it to stop building models whose reasoning humans can no longer audit, according to a Wall Street Journal report. The three, Jasmine Wang, Tomek Korbak and Mikita Balesni, argue that "OpenAI and its industry rivals should stop pursuing the development of models whose reasoning is difficult to monitor and audit."

Korbak and Balesni were lead authors of the "Chain of Thought Monitorability" paper, which makes the technical case for keeping frontier models' step-by-step reasoning legible to human reviewers. Korbak served as OpenAI's technical contact for the METR organization, the external evaluator, and Balesni helped the board itself draft an industry safety pledge. The three use those sanctioned roles to argue they never exceeded their professional duties.

The letter, sent to the board on Wednesday, says the firings are creating a "chilling" atmosphere for employees who remain at the firm.

OpenAI's account differs. The company says an internal investigation found the three mishandled sensitive data, including infrastructure architecture, and has denied the terminations were related to the employees raising safety concerns. The dispute lands in the middle of a dense run of OpenAI coverage we've tracked this quarter, and now reaches the people who co-authored the industry's main playbook for keeping model reasoning auditable.