Saphra and Wiegreffe map four uses of 'mechanistic' in AI
TL;DR
- Two researchers identify four separate uses of 'mechanistic' in interpretability, from a strict causality requirement to any exploration of a model's internals.
- Saphra and Wiegreffe trace the term to a distinct research community that grew in parallel with the older NLP interpretability field.
- They argue the traditional NLP interpretability community has 'come to embrace' the broad cultural version of the label, folding into the newer name.
Two interpretability researchers argue the field can't agree what 'mechanistic' means, and the confusion is not just jargon.
In a position paper accepted at the BlackBoxNLP workshop at EMNLP 2024, Naomi Saphra and Sarah Wiegreffe count four separate uses of the word. The narrowest technical version 'requires a claim of causality.' A broader technical one 'allows for any exploration of a model's internals.' Then there is a narrow cultural definition describing a distinct movement, and a broad cultural definition covering the whole interpretability field.
'We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community,' the authors write. To trace it, they lay out a history of NLP interpretability and the formation of the separate, parallel 'mechanistic' community that grew alongside it.
Their read on where things landed: the traditional NLP interpretability community has 'come to embrace' the broad cultural definition, absorbing itself into the newer label.
Shared on Bluesky by 1 AI expert
Originally reported by arxiv.org
Read the original article →Original headline: Mechanistic?