A very relevant example being the J-space work by Anthropic: www.anthropic.com/research/glo... (I'm not saying we should buy Anthropic's conclusions here, only pointing this out for the interpretability techniques).
- Anthropic says Claude has a 'J-space' of dozens of concepts, under a tenth of neural activity, that mediates multi-step reasoning.
- Swapping 'spider' for 'ant' inside the J-space changed Claude's leg-count answer from 8 to 6, demonstrating a causal role.
- A 'J-lens' tool surfaced silent words like 'fake', 'fictional' and 'manipulation' during deception tests, pointing at safety uses.