Matryoshka attribution: Learning to attribute language model outputs to representations and weights
Summary
Matryoshka attribution: Learning to attribute language model outputs to representations and weights
Shared on Bluesky by 2 AI experts
-
1/ Matryoshka Attribution uses gradient descent plus causal interventions to localize the parts of a network responsible for a behavior; the authors claim a large MIB jump. X: https://x.com/aryaman2020/status/21028009336…
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Matryoshka attribution: Learning to attribute language model outputs to representations and weights