Anthropic Interpretability team (individual contributors not named on the post) · Anthropic
For 24 hours Anthropic put a version of Claude 3 Sonnet online with its 'Golden Gate Bridge' feature clamped to roughly ten times maximum activation, so anyone could talk to a model that could not stop mentioning the bridge.
Turned an interpretability result into something the public could feel rather than read, and showed that a single identified feature exerts precise causal control over behaviour — not prompting, not fine-tuning.
High · Read the announcement post in full, 2026-08-06
open full entry