INTERP ATLAS

visual form

Scatter plot

3 entries

2026 (1)

IA-005

HeadVis is showing , about , and .

R. Luger; Harish Kamath; Doug Finkbeiner; Purvi Goel; Adam Jermyn; Sam Zimmerman; Joshua Batson; Tom Conerly · Anthropic

An interactive tool for interrogating attention heads, released as the research output in its own right. Its central finding is methodological: a head's behaviour on the full data distribution rarely matches what a narrow task suggests.

Reframes the tool as the contribution, and supplies the concrete examples the field needs to attack attention decomposition.

HighRead the full article, 2026-08-06

open full entry

2023 (1)

IA-010

AttentionViz is showing , about .

Catherine Yeh; Yida Chen; Aoyu Wu; Cynthia Chen; Fernanda Viégas; Martin Wattenberg · Harvard University

Visualises a joint embedding of the query and key vectors a transformer uses to compute attention, which makes it possible to see global patterns across many input sequences rather than one prompt at a time.

Shifted attention visualisation from single-example to distribution-level — the exact move Anthropic's HeadVis names as its closest prior work three years later.

HighRead the arXiv abstract page in full, 2026-08-06. The live demo at attentionviz.com was not inspected — client-rendered, fetch returned an empty page shell.

open full entry

2019 (1)

IA-007

Activation Atlas is and showing , about .

Activation Atlas
2026-08-12

Shan Carter; Zan Armstrong; Ludwig Schubert; Ian Johnson; Chris Olah · Google Brain; OpenAI

Renders millions of activations from an image classifier as feature-inversion images laid out on a single navigable map, so you can pan across the concepts a network has learned the way you would read an atlas.

Made a model's whole learned concept space visible at once rather than one neuron at a time — the clearest ancestor of the feature-neighbourhood maps in Scaling Monosemanticity.

HighRead the Distill article and its citation metadata, 2026-08-06

open full entry