transformer-circuits.pub web signal

Anthropic team characterizes interference weights in a tiny LM

TL;DR

  • Anthropic researchers Nicholas L. Turner, Jeffrey Wu and Joshua Batson define 'interference weights' as residual-stream interactions that are irrelevant or harmful to model behavior.
  • The stated hope is that removing them recovers a sparse model whose remaining weights reflect 'the circuits the model actually hoped to learn.'
  • The work extends a July 2025 informal note by Chris Olah, Turner and Tom Conerly that framed interference weights as a bridge to global circuit analysis.

Anthropic's interpretability team has extended its 2025 toy-model work on "interference weights" into a real, if small, language model. In Characterizing interference weights in a tiny language model, Nicholas L. Turner, Jeffrey Wu and Joshua Batson define the target as "linear interactions of interpretable model components through the low-dimensional residual stream which are either irrelevant or harmful to the model's behavior."

The setup addresses a familiar bottleneck in mechanistic interpretability: even when researchers recover clean features, the virtual weights between those features are cluttered by connections that reflect superposition rather than intended computation. "If we could identify interference weights accurately," the authors write, "we could hope to remove them and recover a sparse model whose remaining weights reflect the circuits the model actually hoped to learn."

The direction picks up from an informal July 2025 note by Chris Olah, Turner and Tom Conerly, which pitched interference weights as the missing step between per-example attribution graphs and global circuit analysis. Two of the researchers we track shared the post the day it went live, and two also passed along this follow-up paper.

Shared on Bluesky by 2 AI experts