Peng, Kleinberg, Garg: large SAEs recover features in order
TL;DR
- The paper's 'recovery principle' predicts sparse autoencoders of increasing size recover an increasing prefix of the most prevalent atoms in the training data.
- Three predictions were tested on SAEs ranging in size from 512 to 131,072, trained on two large embedding models.
- The authors argue the results contradict conventional wisdom that SAE features are unstable and 'split' as size increases.
Sparse autoencoders behave more predictably at scale than interpretability researchers had assumed, according to a new theory and empirical tests from Kenny Peng, Jon Kleinberg and Nikhil Garg. In their arxiv paper, the three propose a theory of 'atomic features' and derive what they call a 'recovery principle': as sparse dictionaries like SAEs grow larger, they recover 'an increasing prefix of the most prevalent atoms in the training data.'
The principle yields three testable predictions, which the paper then checks against SAEs ranging in size from 512 to 131,072, trained on two large embedding models. The predictions are that features in small SAEs persist in larger ones, that SAEs trained on different data share features prevalent in both, and that 'sufficiently large SAEs recover both parent and child features.'
All three held on the authors' tests.
That framing is a deliberate push against a common view in the field. The paper names 'conventional wisdom that SAE features are unstable and "split" as size increases'; its results argue instead that growth is additive. The authors conclude the findings 'suggest the promise of a scientific theory of representations based on atomic features' and, practically, 'the promise of scaling SAEs.' The empirical work was done on embedding models rather than decoder-only language models. Two interpretability researchers we track had the paper circulating the day it posted.
Shared on Bluesky by 2 AI experts
-
Very excited to have this out -- @kennypeng.bsky.social 's brainchild in thinking about the implications of sparse representations. He's on the job market, and watch out for some exciting tools we're releasing soon! arx…
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: A Testable Theory of Atomic Features