INTERP ATLAS

method

activation-patching

1 entry

2025 (1)

IA-004

On the Biology of a Large Language Model is showing , about , , and .

Jack Lindsey; Emmanuel Ameisen; Adam Pearce; Joshua Batson; et al. · Anthropic

Ten case studies of Claude 3.5 Haiku's internal mechanisms rendered as explorable attribution graphs — showing it plans rhymes ahead, reasons across a shared multilingual concept space, and sometimes fabricates reasoning backwards from a hinted answer.

The clearest existing demonstration that a model's actual reasoning can diverge from its stated reasoning, made legible by graph.

HighRead the Anthropic research write-up in full, 2026-08-06

open full entry