INTERP ATLAS

modality

language

9 entries

2026 (1)

IA-005

HeadVis is showing , about , and .

R. Luger; Harish Kamath; Doug Finkbeiner; Purvi Goel; Adam Jermyn; Sam Zimmerman; Joshua Batson; Tom Conerly · Anthropic

An interactive tool for interrogating attention heads, released as the research output in its own right. Its central finding is methodological: a head's behaviour on the full data distribution rarely matches what a narrow task suggests.

Reframes the tool as the contribution, and supplies the concrete examples the field needs to attack attention decomposition.

HighRead the full article, 2026-08-06

open full entry

2025 (1)

IA-004

On the Biology of a Large Language Model is showing , about , , and .

Jack Lindsey; Emmanuel Ameisen; Adam Pearce; Joshua Batson; et al. · Anthropic

Ten case studies of Claude 3.5 Haiku's internal mechanisms rendered as explorable attribution graphs — showing it plans rhymes ahead, reasons across a shared multilingual concept space, and sometimes fabricates reasoning backwards from a hinted answer.

The clearest existing demonstration that a model's actual reasoning can diverge from its stated reasoning, made legible by graph.

HighRead the Anthropic research write-up in full, 2026-08-06

open full entry

2024 (4)

IA-009

Transformer Explainer is showing , about .

Aeree Cho; Grace C. Kim; Alexander Karpekov; Alec Helbling; Zijie J. Wang; Seongmin Lee; Benjamin Hoover; Duen Horng (Polo) Chau · Georgia Institute of Technology (Polo Club of Data Science)

An interactive tool that runs a live GPT-2 instance in your browser and visualises every stage of the forward pass — embeddings, attention, MLP, output probabilities — updating as you type your own text.

The most accessible route into transformer internals that exists: no installation, no GPU, no prior knowledge. That is a different kind of contribution from any research result in this ledger, and arguably a wider-reaching one.

HighRead the arXiv abstract page in full, 2026-08-06. The live tool itself was not inspected — it is client-rendered and the fetch returned an empty page shell.

open full entry
IA-011

TalkTuner is showing , about and .

Yida Chen; Aoyu Wu; Trevor DePodesta; Catherine Yeh; Kenneth Li; Nicholas Castillo Marin; Oam Patel; Jan Riecke; Shivam Raval; Olivia Seow; Martin Wattenberg; Fernanda Viégas · Harvard University

A dashboard that sits beside a chatbot and shows, in real time, what the model has internally inferred about the user's age, gender, education and socioeconomic status — and lets the user edit those inferences and watch the responses change.

Almost everything else in this ledger is built for researchers; this is built for the person being modelled, and a user study found it helped participants expose the system's biased behaviour.

HighRead the arXiv abstract page in full, 2026-08-06. Project page (bit.ly/talktuner-project-page) not fetched.

open full entry
IA-008

Golden Gate Claude has no visual form and shows , about .

Anthropic Interpretability team (individual contributors not named on the post) · Anthropic

For 24 hours Anthropic put a version of Claude 3 Sonnet online with its 'Golden Gate Bridge' feature clamped to roughly ten times maximum activation, so anyone could talk to a model that could not stop mentioning the bridge.

Turned an interpretability result into something the public could feel rather than read, and showed that a single identified feature exerts precise causal control over behaviour — not prompting, not fine-tuning.

HighRead the announcement post in full, 2026-08-06

open full entry
IA-003

Scaling Monosemanticity is showing , about and .

Adly Templeton; Tom Conerly; Jonathan Marcus; Jack Lindsey; Trenton Bricken; Brian Chen; Adam Jermyn; et al. · Anthropic

First demonstration that sparse autoencoders scale from toy models to a frontier production model, published as a browsable index of millions of features — including the Golden Gate Bridge feature that later became a public demo.

The moment interpretability stopped being a toy-model science, and the origin of the field's most famous public artefact.

HighRead the article; cross-referenced in BlueDot and ACX pieces, 2026-08-06

open full entry

2023 (3)

IA-002

Towards Monosemanticity is showing , about , and .

Trenton Bricken; Adly Templeton; Joshua Batson; Brian Chen; Adam Jermyn; Tom Conerly; Nicholas L. Turner; Cem Anil; Carson Denison; Amanda Askell; et al. · Anthropic

The paper that made sparse autoencoders the field's dominant method, published with a browsable interface over every extracted feature so readers could check the monosemanticity claim themselves rather than take the authors' word for it.

Turned a contested claim (features are more interpretable than neurons) into something a reader could audit by clicking.

HighRead the article and its setup/interface section, 2026-08-06

open full entry
IA-010

AttentionViz is showing , about .

Catherine Yeh; Yida Chen; Aoyu Wu; Cynthia Chen; Fernanda Viégas; Martin Wattenberg · Harvard University

Visualises a joint embedding of the query and key vectors a transformer uses to compute attention, which makes it possible to see global patterns across many input sequences rather than one prompt at a time.

Shifted attention visualisation from single-example to distribution-level — the exact move Anthropic's HeadVis names as its closest prior work three years later.

HighRead the arXiv abstract page in full, 2026-08-06. The live demo at attentionviz.com was not inspected — client-rendered, fetch returned an empty page shell.

open full entry
IA-006

Neuronpedia is and showing and , about .

Neuronpedia
2026-08-12

Johnny Lin (creator), with community contributors · Neuronpedia / Decode Research

The field's central public platform. It hosts feature dashboards, attribution graphs, steering and demos across dozens of open models — including the official interactive releases for Anthropic's Jacobian Lens, Natural Language Autoencoders, Assistant Axis and Circuit Tracer, and DeepMind's Gemma Scope.

The closest thing interpretability has to a public commons, and the reason a frontier-lab result can now be poked at by an outsider the week it ships.

HighRead the neuronpedia.org homepage in full, 2026-08-06

open full entry