IA-006 / 2023
Neuronpedia is a dashboardA single screen putting several different readouts side by side.Several visual forms arranged together as one interface, read as a unit.glossary and a node-link graphDots joined by lines, showing what connects to what.Nodes joined by edges, showing what connects to or causes what.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary and circuitsA chain of parts inside the model that work together to do one job.Connected components and the paths between them.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.

- ID
- IA-006
- Name
- Neuronpedia
- Artefact type
- Platform
- Visual form
- DashboardA single screen putting several different readouts side by side.Several visual forms arranged together as one interface, read as a unit.glossaryNode-link graphDots joined by lines, showing what connects to what.Nodes joined by edges, showing what connects to or causes what.glossary
- Makes visible
- FeaturesA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossaryCircuitsA chain of parts inside the model that work together to do one job.Connected components and the paths between them.glossary
- Role of the visual
- Instrument
- Runnable by a visitor
- On your own inputs
- Date published
- 2023
- Year
- 2023
- Link status
- Live
- Built by (people)
- Johnny Lin (creator), with community contributors
- Organisation
- Neuronpedia / Decode Research
- Venue / published in
- Self-published (web)
- Sector
- Non-profit / community
- Country / region
- US
- Open source
- Yes
- Method / technique
- Hosts SAE and transcoder feature dashboards, attribution graphs (Circuit Tracer), activation steering, probes, custom vectors, and lab-released lenses (Jacobian Lens, Natural Language Autoencoders, Assistant Axis)
- Model(s) studied
- GPT-2 Small, Pythia-70M-deduped, Gemma 2/3/4, Qwen 3/3.5/3.6, Llama 3.1/3.3, Olmo 3, GPT-OSS-20B and others
- Access required
- Open weights
- What it visualises
- Individual latents with top activations, top logits and activation density; attribution graphs on custom prompts; steering effects; over five terabytes of activations, explanations and metadata
- Interaction affordances
- Search 50M+ latents by explanation or by running text through a model; browse per-feature dashboards with permanent URLs; compile shareable lists; embed as iframe; steer with adjustable strength, temperature and seed; trace circuits on your own prompts; full API
- Pragmatic vs Basic science
- Both
- Reverse-eng vs Concept-based
- Both
- Observational vs Interventional
- Both
- Intended audience
- Mixed
- Description (card)
- The field's central public platform. It hosts feature dashboards, attribution graphs, steering and demos across dozens of open models — including the official interactive releases for Anthropic's Jacobian Lens, Natural Language Autoencoders, Assistant Axis and Circuit Tracer, and DeepMind's Gemma Scope.
- Why it matters
- The closest thing interpretability has to a public commons, and the reason a frontier-lab result can now be poked at by an outsider the week it ships.
- Visual / design notes
- One dashboard per feature, each with a permanent URL and iframe embedding. That single decision is what let it become shared infrastructure rather than one lab's internal tool.
- Tags
- modality:languageText. Models that read and write words.Models that generate or process text.glossarymodality:generalNot tied to one kind of data. Applies across different models and media.Not tied to any particular model or data type.glossarymethod:sparse-autoencoderA technique for pulling a model's tangled internals apart into separate, nameable pieces.Learning an overcomplete, sparsely activating basis for a layer's activations.glossarymethod:transcoderA variant that reconstructs what a layer passes onward, rather than what it received.A sparse dictionary trained to reproduce a layer's output rather than its input.glossarymethod:attribution-graphA map of which internal pieces caused which, for one specific prompt.A causal graph of interpretable components explaining a single forward pass.glossarymethod:steeringNudging the model's internals mid-thought to change what it says.Adding or subtracting a direction in activation space in order to alter behaviour.glossaryphenomenon:feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary
- Media files
- IA-006-neuronpedia-circuit-tracer--2026-08-12.jpg
- Citation
- Lin, J., 2023. Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks. Software available from neuronpedia.org
- Related entries
- IA-003; IA-004; IA-005
- Confidence
- High
- Source of info
- Read the neuronpedia.org homepage in full, 2026-08-06
- Date added
- 2026-08-06
- Added by
- Claude
- Notes
- Supported by Decode Research, Open Philanthropy, the Long Term Future Fund, AISTOF, Anthropic and Manifund. Hosts releases from Anthropic, DeepMind, OpenAI, EleutherAI, Apollo, Fudan OpenMOSS, Goodfire and Stanford NLP. Decision pending: whether individual hosted demos get their own rows.