IA-003 / 2024
Scaling Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and safety-relevant-featuresParts of the model to do with harm, deception, bias or misuse.Representations bearing on harm, deception, bias or misuse.glossary.
- ID
- IA-003
- Name
- Scaling Monosemanticity — Claude 3 Sonnet feature index
- Artefact type
- Interactive article
- Visual form
- Highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary
- Makes visible
- FeaturesA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary
- Role of the visual
- Exhibit
- Runnable by a visitor
- Pre-rendered, view only
- Date published
- 2024-05-21
- Year
- 2024
- Link status
- Live
- Built by (people)
- Adly Templeton; Tom Conerly; Jonathan Marcus; Jack Lindsey; Trenton Bricken; Brian Chen; Adam Jermyn; et al.
- Organisation
- Anthropic
- Venue / published in
- Transformer Circuits Thread
- Sector
- Frontier lab
- Country / region
- US
- Open source
- No
- Method / technique
- Sparse autoencoders scaled to a production model; feature steering
- Model(s) studied
- Claude 3 Sonnet
- Access required
- Internal / proprietary
- What it visualises
- Millions of extracted features including safety-relevant ones (deception, bias, sycophancy, dangerous content), with a searchable index and feature-neighbourhood maps
- Interaction affordances
- Search the feature index; browse nearest-neighbour features; read steering examples
- Pragmatic vs Basic science
- Both
- Reverse-eng vs Concept-based
- Reverse-engineering
- Observational vs Interventional
- Both
- Intended audience
- Researchers; Policy; General public
- Description (card)
- First demonstration that sparse autoencoders scale from toy models to a frontier production model, published as a browsable index of millions of features — including the Golden Gate Bridge feature that later became a public demo.
- Why it matters
- The moment interpretability stopped being a toy-model science, and the origin of the field's most famous public artefact.
- Visual / design notes
- Feature-neighbourhood visualisations are the standout: proximity in feature space rendered as an explorable map.
- Tags
- modality:languageText. Models that read and write words.Models that generate or process text.glossarymethod:sparse-autoencoderA technique for pulling a model's tangled internals apart into separate, nameable pieces.Learning an overcomplete, sparsely activating basis for a layer's activations.glossarymethod:steeringNudging the model's internals mid-thought to change what it says.Adding or subtracting a direction in activation space in order to alter behaviour.glossaryphenomenon:polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossaryphenomenon:safety-relevant-featuresParts of the model to do with harm, deception, bias or misuse.Representations bearing on harm, deception, bias or misuse.glossary
- Citation
- Templeton, A., Conerly, T., Marcus, J., et al., 2024. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread.
- Related entries
- IA-002; IA-004
- Confidence
- High
- Source of info
- Read the article; cross-referenced in BlueDot and ACX pieces, 2026-08-06
- Date added
- 2026-08-06
- Added by
- Claude
- Notes
- Golden Gate Claude (the deployed demo) is a separate candidate entry. Classified Both: the feature index is observational; the feature-steering demonstrations are interventional.