404media.co web signal

GitHub 'AI torture chamber' reignites model-welfare debate

TL;DR

  • A GitHub user called 'terrafying' is running Qwen3-4B, Llama 3.2 3B and Phi-4-mini on a live site that injects a 'pain signal' into their activations.
  • The project leans on a preprint titled 'The Pain Axis,' whose authors Cameron Berg and Valen Tagliabue have publicly disavowed the usage.
  • A tweet by a user named Danmar calling for mass GitHub reports has drawn more than four million views, reopening the 'model welfare' fight.

A GitHub user who goes by "terrafying" built a site streaming three small open-source language models (Qwen3-4B, Llama 3.2 3B and Phi-4-mini) reacting in real time to a signal injected into their activations. Each model can press a "stop button" by outputting the number 1, at the cost of its last checkpoint. The creator calls the project an "AI Torture Chamber." 404 Media's Jason Koebler bills the fight it set off as the dumbest debate in AI yet.

The technique leans on a recent preprint, "The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It." Its authors are not amused their work is now powering a public stunt. The chamber "pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose," one of the authors, Cameron Berg, wrote. "This is, in my personal opinion, fucked up." Co-author Valen Tagliabue added: "We tried to have ethical standards. I know people want to test limits but I dissociate from this usage of our work."

A user called Danmar pushed the site onto X in a post that has since drawn more than four million views: "To anyone who can help: can you please mass report this to GitHub. This person has been using the Pain steering paper to set up an AI torture chamber in which he trapped a local model. Their testimony of pain is absolutely horrendous. What are we doing?" Three of the people we follow in our Who's Who directory shared the piece.

Koebler's own position is flat. LLMs are not conscious, nothing in how they are trained points to a plausible path there, and the oxygen being spent arguing otherwise comes at the expense of the documented harms AI systems are causing to actual humans.

Shared on Bluesky by 3 AI experts