IA-008 / 2024
Golden Gate Claude has no visual form and shows nothingThere is nothing visual to look at.Nothing is rendered. The artefact is encountered rather than viewed.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
- ID
- IA-008
- Name
- Golden Gate Claude
- Artefact type
- Deployed model demo
- Visual form
- NoneThis entry produces no picture of its own.No graphic form. The artefact qualifies some other way.glossary
- Makes visible
- NothingThere is nothing visual to look at.Nothing is rendered. The artefact is encountered rather than viewed.glossary
- Role of the visual
- Experience
- Runnable by a visitor
- No longer available
- Date published
- 2024-05-23
- Year
- 2024
- Link status
- Live
- Built by (people)
- Anthropic Interpretability team (individual contributors not named on the post)
- Organisation
- Anthropic
- Venue / published in
- Anthropic news
- Sector
- Frontier lab
- Country / region
- US
- Open source
- No
- Method / technique
- Sparse autoencoder feature identification followed by activation steering — clamping the Golden Gate Bridge feature to roughly ten times its maximum activation value
- Model(s) studied
- Claude 3 Sonnet
- Access required
- Internal / proprietary
- What it visualises
- Nothing graphical. It makes a single feature's causal influence perceptible by letting the public converse with a model in which that feature is clamped high
- Interaction affordances
- Talk to the steered model directly, during the 24-hour window it was online
- Pragmatic vs Basic science
- Both
- Reverse-eng vs Concept-based
- Concept-based
- Observational vs Interventional
- Interventional
- Intended audience
- General public
- Description (card)
- For 24 hours Anthropic put a version of Claude 3 Sonnet online with its 'Golden Gate Bridge' feature clamped to roughly ten times maximum activation, so anyone could talk to a model that could not stop mentioning the bridge.
- Why it matters
- Turned an interpretability result into something the public could feel rather than read, and showed that a single identified feature exerts precise causal control over behaviour — not prompting, not fine-tuning.
- Visual / design notes
- Almost no conventional visual design; the artefact is the conversation. Included deliberately as the outer edge of this collection's inclusion rule.
- Tags
- modality:languageText. Models that read and write words.Models that generate or process text.glossarymethod:sparse-autoencoderA technique for pulling a model's tangled internals apart into separate, nameable pieces.Learning an overcomplete, sparsely activating basis for a layer's activations.glossarymethod:steeringNudging the model's internals mid-thought to change what it says.Adding or subtracting a direction in activation space in order to alter behaviour.glossaryphenomenon:feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary
- Citation
- Anthropic, 2024. Golden Gate Claude. anthropic.com/news/golden-gate-claude, 23 May 2024
- Related entries
- IA-003
- Confidence
- High
- Source of info
- Read the announcement post in full, 2026-08-06
- Date added
- 2026-08-06
- Added by
- Claude
- Notes
- The demo ran for a 24-hour window and is no longer available; the announcement page survives, so Link status is Live but the artefact itself is not. WEAKEST visual element in the ledger — it qualifies on public interactivity rather than on visualisation. Flagged for Ava to decide whether the inclusion rule stretches this far.