Anthropic ships Golden Gate Claude, an interpretability demo
TL;DR
- Anthropic released Golden Gate Claude, a Claude 3 Sonnet variant with one internal feature amplified so the Golden Gate Bridge intrudes into almost every answer.
- The demo sits on top of Anthropic's paper 'Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet,' which uses dictionary learning to isolate concept features.
- The same technique surfaced features tied to deception, bombs and bioweapons, and racism, hinting at future monitoring or dampening interventions.
Anthropic did something unusual for a research shop this weekend, they shipped an intentionally broken model. For roughly twenty-four hours, users could open a variant of Claude 3 Sonnet, click a bridge icon in the Claude interface, and talk to Golden Gate Claude, a version with one internal feature turned way up so that no matter what you asked, the Golden Gate Bridge found a way into the answer. Simon Willison's writeup called the whole thing "absurdly fun and weird," and the examples earn it: a suggested pelican name of "Golden Gate - This iconic bridge name would be a fitting moniker," and a chocolate-pretzel recipe that drifts into "Golden Gate Bridge baked bricks" and helicopters.
The mechanic underneath is what makes it more than a party trick. It sits on top of Anthropic's paper, "Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet," which uses a technique called dictionary learning to locate specific combinations of neurons that fire when a particular concept shows up in the input. The bridge is one of those. For the demo, Anthropic amplified that feature's weight during inference. As Willison put it, no matter what question you ask, the Golden Gate Bridge is likely to be involved in the answer in some way.
The reason a novelty release is worth paying attention to is what else the same microscope found. Beyond the playful demo, the Anthropic team reportedly identified features associated with deception, the creation of bombs and bioweapons, and racism. If those features can be located reliably, they can in principle be monitored or dampened rather than only trained around after the fact. Golden Gate Claude is the vaudeville version of a much more serious claim, that you can grab a specific concept inside a frontier model and turn a knob on it.
The honest caveats are what the reporting doesn't give you. There's no public number on how expensive dictionary learning is against frontier-scale models, no evidence yet that a feature like "deception" is as cleanly isolable as a physical landmark, and no sense of how much real behavior you can shape by tuning a handful of features versus needing thousands. And a boosted model that suggests seawater and helicopters is not fit for anything, which is why the demo came down inside a day.
What's interesting is the direction. If steering by feature ever becomes cheap and precise, safety teams get an intervention that lives below the prompt and fine-tune layers, and product teams eventually get personality dials that don't require a fresh training run. Take the specifics as reported, not settled, but this is the part of interpretability work that starts to look like a product surface rather than a paper.
Shared on Bluesky by 2 AI experts
-
Golden Gate Claude was actually the peak of human civilization and it has all been downhill since then simonwillison.net/2024/May/24/...
View on Bluesky →
Originally reported by simonwillison.net
Read the original article →Original headline: Golden Gate Claude