INTERP ATLAS

phenomenon

polysemanticity

3 entries

2026 (1)

IA-005

HeadVis is showing , about , and .

R. Luger; Harish Kamath; Doug Finkbeiner; Purvi Goel; Adam Jermyn; Sam Zimmerman; Joshua Batson; Tom Conerly · Anthropic

An interactive tool for interrogating attention heads, released as the research output in its own right. Its central finding is methodological: a head's behaviour on the full data distribution rarely matches what a narrow task suggests.

Reframes the tool as the contribution, and supplies the concrete examples the field needs to attack attention decomposition.

HighRead the full article, 2026-08-06

open full entry

2024 (1)

IA-003

Scaling Monosemanticity is showing , about and .

Adly Templeton; Tom Conerly; Jonathan Marcus; Jack Lindsey; Trenton Bricken; Brian Chen; Adam Jermyn; et al. · Anthropic

First demonstration that sparse autoencoders scale from toy models to a frontier production model, published as a browsable index of millions of features — including the Golden Gate Bridge feature that later became a public demo.

The moment interpretability stopped being a toy-model science, and the origin of the field's most famous public artefact.

HighRead the article; cross-referenced in BlueDot and ACX pieces, 2026-08-06

open full entry

2023 (1)

IA-002

Towards Monosemanticity is showing , about , and .

Trenton Bricken; Adly Templeton; Joshua Batson; Brian Chen; Adam Jermyn; Tom Conerly; Nicholas L. Turner; Cem Anil; Carson Denison; Amanda Askell; et al. · Anthropic

The paper that made sparse autoencoders the field's dominant method, published with a browsable interface over every extracted feature so readers could check the monosemanticity claim themselves rather than take the authors' word for it.

Turned a contested claim (features are more interpretable than neurons) into something a reader could audit by clicking.

HighRead the article and its setup/interface section, 2026-08-06

open full entry