INTERP ATLAS

phenomenon

feature-interpretation

5 entries

2024 (1)

IA-008

Golden Gate Claude has no visual form and shows , about .

Anthropic Interpretability team (individual contributors not named on the post) · Anthropic

For 24 hours Anthropic put a version of Claude 3 Sonnet online with its 'Golden Gate Bridge' feature clamped to roughly ten times maximum activation, so anyone could talk to a model that could not stop mentioning the bridge.

Turned an interpretability result into something the public could feel rather than read, and showed that a single identified feature exerts precise causal control over behaviour — not prompting, not fine-tuning.

HighRead the announcement post in full, 2026-08-06

open full entry

2023 (2)

IA-002

Towards Monosemanticity is showing , about , and .

Trenton Bricken; Adly Templeton; Joshua Batson; Brian Chen; Adam Jermyn; Tom Conerly; Nicholas L. Turner; Cem Anil; Carson Denison; Amanda Askell; et al. · Anthropic

The paper that made sparse autoencoders the field's dominant method, published with a browsable interface over every extracted feature so readers could check the monosemanticity claim themselves rather than take the authors' word for it.

Turned a contested claim (features are more interpretable than neurons) into something a reader could audit by clicking.

HighRead the article and its setup/interface section, 2026-08-06

open full entry
IA-006

Neuronpedia is and showing and , about .

Neuronpedia
2026-08-12

Johnny Lin (creator), with community contributors · Neuronpedia / Decode Research

The field's central public platform. It hosts feature dashboards, attribution graphs, steering and demos across dozens of open models — including the official interactive releases for Anthropic's Jacobian Lens, Natural Language Autoencoders, Assistant Axis and Circuit Tracer, and DeepMind's Gemma Scope.

The closest thing interpretability has to a public commons, and the reason a frontier-lab result can now be poked at by an outsider the week it ships.

HighRead the neuronpedia.org homepage in full, 2026-08-06

open full entry

2019 (1)

IA-007

Activation Atlas is and showing , about .

Activation Atlas
2026-08-12

Shan Carter; Zan Armstrong; Ludwig Schubert; Ian Johnson; Chris Olah · Google Brain; OpenAI

Renders millions of activations from an image classifier as feature-inversion images laid out on a single navigable map, so you can pan across the concepts a network has learned the way you would read an atlas.

Made a model's whole learned concept space visible at once rather than one neuron at a time — the clearest ancestor of the feature-neighbourhood maps in Scaling Monosemanticity.

HighRead the Distill article and its citation metadata, 2026-08-06

open full entry

2018 (1)

IA-001

The Building Blocks of Interpretability is and showing and , about .

Chris Olah; Arvind Satyanarayan; Ian Johnson; Shan Carter; Ludwig Schubert; Katherine Ye; Alexander Mordvintsev · Google Brain

The founding text for interactive interpretability. Argues that interpretability techniques studied in isolation are far weaker than the interfaces you get by composing them, and demonstrates this with a set of live, hoverable interfaces over an image classifier.

Established that the interface IS the contribution — the template every entry in this ledger inherits from.

HighRead in full (article + Distill metadata), 2026-08-06

open full entry