INTERP ATLAS

method

attribution-graph

2 entries

2025 (1)

IA-004

On the Biology of a Large Language Model is showing , about , , and .

Jack Lindsey; Emmanuel Ameisen; Adam Pearce; Joshua Batson; et al. · Anthropic

Ten case studies of Claude 3.5 Haiku's internal mechanisms rendered as explorable attribution graphs — showing it plans rhymes ahead, reasons across a shared multilingual concept space, and sometimes fabricates reasoning backwards from a hinted answer.

The clearest existing demonstration that a model's actual reasoning can diverge from its stated reasoning, made legible by graph.

HighRead the Anthropic research write-up in full, 2026-08-06

open full entry

2023 (1)

IA-006

Neuronpedia is and showing and , about .

Neuronpedia
2026-08-12

Johnny Lin (creator), with community contributors · Neuronpedia / Decode Research

The field's central public platform. It hosts feature dashboards, attribution graphs, steering and demos across dozens of open models — including the official interactive releases for Anthropic's Jacobian Lens, Natural Language Autoencoders, Assistant Axis and Circuit Tracer, and DeepMind's Gemma Scope.

The closest thing interpretability has to a public commons, and the reason a frontier-lab result can now be poked at by an outsider the week it ships.

HighRead the neuronpedia.org homepage in full, 2026-08-06

open full entry