Glossary

How this index works

Academia

A university or academic research group.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Access required

The level of model access a person would need to reproduce the work.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

API

Requires gated or paid inference access, but not weights.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Archived

Preserved in static or historical form and no longer maintained.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Art object / installation

Work made primarily for aesthetic or critical reception rather than scientific reporting.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Artefact type

The form the work takes, independent of what it contains or argues.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Artist

A practitioner working primarily within an art context.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Atlas schema

Coined for this index to describe its own structure. Not a claim about the field, and no external source exists.

audio

Models that process sound or speech.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Ava's framing

Supplied by Ava in conversation rather than drawn from a source.

Basic science

Work seeking to reverse-engineer the model's general structure, irrespective of any single behaviour.

Ava's framing · Ava, in conversation, 2026-08-06. Consonant with Nanda et al., 2025, 'A Pragmatic Vision for Interpretability'

Both

Work pursuing a specific behaviour in order to establish general structure, or the reverse.

Ava's framing · Ava, in conversation, 2026-08-06. Consonant with Nanda et al., 2025, 'A Pragmatic Vision for Interpretability'

Browsable index / database

A searchable collection of results presented for inspection.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Built by (people)

The named individual contributors, in the order the source credits them.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Citation

A full bibliographic reference in a consistent style.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Claude's coinage

The practice is real but the LABEL was invented by Claude and is not established. Rename before publishing where a standard term exists.

code

Models whose subject matter is programming languages.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Code repo

A link to that source code.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Common ML usage

In wide use across machine learning with no single coining source. Safe to use, not attributable.

Confidence

How well evidenced this row is, judged by what the person writing it actually read.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Controlled value

A value that may only be chosen from the list on the Vocabularies sheet, never typed freely.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Country / region

Where the producing institution is based.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Dashboard

A live readout of a system's internal state, displayed alongside that system in use.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Date added / Added by

When the row was created and by whom, human or model.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Date published

The date the artefact first became publicly available, in ISO form where known.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Dead

The URL no longer resolves to the artefact.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Deployed model demo

A model made publicly usable in a deliberately modified state, as a demonstration.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Description (card)

One or two sentences describing the artefact to a stranger, written to be published rather than as a note to self.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Exhibit

The visual exists so that a reader can inspect and audit a claim which is itself stated in prose.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava. Not a field taxonomy — it does not exist elsewhere

Experience

The artefact is undergone rather than read, and its value lies in what it does to the visitor.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava. Not a field taxonomy — it does not exist elsewhere

Explanation

The visual communicates a result established elsewhere, serving teaching rather than discovery.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava. Not a field taxonomy — it does not exist elsewhere

Field term

Established vocabulary in the published interpretability literature. The citation is the earliest use Claude could verify, not a guaranteed first use.

Finding

The result is constituted by the visual, such that removing the images would remove the contribution.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava. Not a field taxonomy — it does not exist elsewhere

Free text

A field written in prose, with no fixed list of permitted values.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Frontier lab

An organisation training state-of-the-art general-purpose models.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Frontier lab + academia

Joint authorship spanning both.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

general

Not tied to one kind of data. Applies across different models and media.

Not tied to any particular model or data type.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

General public

Non-specialists, with no assumed technical background.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

genomics

Models of genetic sequence.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

High

The primary source was read in full by whoever wrote the row.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

HOW TO READ THE THREE PROVENANCE COLUMNS

Origin says who coined the TERM. Coined in gives the citation. Definition drafted from says what Claude actually read to write the sentence — which is not the same thing.

ID

A permanent identifier for one artefact, never reused, so that a retired entry's number can never resolve to a different record.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

image-generation

Models that produce images, such as diffusion models.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Independent researcher

A person working outside any institution.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Instrument

Others can operate it on their own inputs or models to produce findings that did not previously exist.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava. Not a field taxonomy — it does not exist elsewhere

Intended audience

Who the artefact was designed to be understood by.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Interaction affordances

What a person can actually do with the artefact, verb by verb.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Interactive article

A written argument in which the figures respond to the reader.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Interactive tool

A standalone application whose purpose is to let a user investigate a model.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Internal / proprietary

Requires access available only within the producing organisation.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

language

Text. Models that read and write words.

Models that generate or process text.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Live

The URL resolves and the artefact functions as described.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Low

The row rests on inference and has not been checked against a source.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Makes visible

The thing rendered on screen. Not what the work is about, which is the phenomenon tag, but what you are actually looking at.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Media files

The image filenames belonging to this entry, semicolon separated.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Medium

The row rests on an abstract, documentation, or a reliable secondary source.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

method

The interpretability technique the work employs.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Method / technique

The interpretability technique or techniques the artefact rests on.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Mixed

Deliberately designed to serve more than one of the above.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

modality

The kind of model or data the work concerns.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Model(s) studied

The specific neural networks the work examines.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

multimodal

Models that process more than one modality jointly.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Name

The artefact's own title, as given by its makers.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Neither / educational

Work that neither diagnoses a behaviour nor uncovers structure, but conveys existing knowledge.

Ava's framing · Ava, in conversation, 2026-08-06. Consonant with Nanda et al., 2025, 'A Pragmatic Vision for Interpretability'

No

No source is published.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

No longer available

The artefact was operable once but is not now.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava

Non-profit / community

An organisation constituted for public benefit rather than profit.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

None

This entry produces no picture of its own.

No graphic form. The artefact qualifies some other way.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

None (pre-rendered)

Requires no model access at all; the outputs are fixed.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Notebook / library

Code published to be run and adapted by other researchers.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Notes

Anything qualifying the row, including unresolved questions and flagged boundary cases.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Nothing

There is nothing visual to look at.

Nothing is rendered. The artefact is encountered rather than viewed.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Observational vs Interventional

Whether the work examines the model unchanged or alters it and observes the consequence.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

On your own inputs

A visitor can supply their own prompt, image or model and see fresh output.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava

Open source

Whether the artefact's own source code is publicly available.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Open weights

Reproducible using publicly downloadable model weights.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Organisation

The institution that produced the work.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Other

A form not covered above; propose a new term rather than reusing this indefinitely.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Partial

Some components are released and others deliberately withheld.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Paywalled

Reachable only behind payment or registration.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

phenomenon

The property or behaviour of the model that the work is about.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Platform

Infrastructure hosting many artefacts, models or datasets originating from multiple parties.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Policy

Regulators, governments and civil society organisations.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Practitioners

People deploying or auditing models professionally.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Pragmatic

Work targeting a specific model behaviour, asking which parts of the network produce it.

Ava's framing · Ava, in conversation, 2026-08-06. Consonant with Nanda et al., 2025, 'A Pragmatic Vision for Interpretability'

Pragmatic vs Basic science

Whether the work targets one specific behaviour or the model's general structure.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Pre-rendered, view only

The visuals are fixed; interaction is limited to navigation such as panning or zooming.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava

Pre-set examples only

A visitor can explore interactively, but only within cases the authors selected in advance.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava

protein

Models of protein sequence or structure.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Purpose

Every column in the Ledger and every permitted value is defined here in one sentence, so that two people filling in the same row reach the same answer.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Researchers

Specialists in interpretability or machine learning.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Reverse-eng vs Concept-based

Whether the work decomposes the network and then names the parts, or proposes concepts and then locates them.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

rl-agent

Agents trained by reinforcement learning to act within an environment.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Role of the visual

The function the visual element performs within the research, which is the central judgement in each row.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Runnable by a visitor

Whether, and how far, a person arriving today can operate the artefact themselves.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

safety-relevant-features

Parts of the model to do with harm, deception, bias or misuse.

Representations bearing on harm, deception, bias or misuse.

Atlas schema · Descriptive. The phrase appears loosely in Templeton et al., 2024 but is not a defined term of art

Sector

The kind of institution that produced the work, recorded so the composition of the field is visible.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Slug

The lowercase hyphenated form of the entry's name, used as its web address.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Source of info

What was read to write this row, and on what date.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Startup

A commercial company whose product is interpretability or model tooling.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Static figures in a paper

Fixed images within a conventional publication.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Students

Learners being taught the material for the first time.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

tabular

Models over structured, columnar data.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Tags

Faceted keywords in facet:value form, drawn only from the controlled facet lists.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Thumbnail URL

An image representing the entry.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Unknown

The role has not been determined.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava. Not a field taxonomy — it does not exist elsewhere

Unknown

Not determined, often because the page could not be loaded.

Atlas schema · Proposed by Claude 2026-08-06, approved by Ava

Unknown

The basis for the row was not recorded.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Unknown

Not yet determined.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Unknown

Not yet determined.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

URL

The canonical public address of the artefact itself, rather than of a paper describing it.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Venue / published in

Where the work appeared, which is often not the same as who produced it.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Video / animation

Time-based linear media.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

vision

Images. Models that look at pictures.

Models that process static images.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Visual / design notes

Observations on the artefact's visual and interaction design, including what makes it work or fail.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Visual form

The shape the picture takes: what you would call the graphic if describing it to someone who could not see it.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

What it visualises

The object being made visible, stated in plain terms a non-specialist could follow.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Why it matters

One line stating the artefact's significance in terms a reader could dispute.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Why no judgements

A tag is a filter, not an opinion; claims about significance belong in Why it matters, where a reader can dispute them.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Why only three

Organisation, artefact type, sector and open source are already columns; tagging them again would create two sources of truth that drift apart.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Year

The publication year held separately, so that entries sort chronologically without parsing dates.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Yes

The artefact's own source code is publicly available under a licence.

Atlas schema · The Interpretability Atlas, 2026 — Ava and Claude

Glossary

Activation grid

Lots of tiny pictures in a grid, each showing what the model saw at that spot.

Many small images laid out in a grid, each one standing for what a part of the model responded to at that position.

Field term · Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K. and Mordvintsev, A., 2018. The Building Blocks of Interpretability. Distill. Verified first-hand 2026-08-22

activation-patching

Swapping one internal value for another to see whether it mattered.

Substituting one activation for another to measure its causal effect on the output.

Field term · Vig et al., 2020; Geiger et al., 2020; named 'activation patching' by Nanda, 2023 (as cited in Sharkey et al., 2025 §2.1.3b)

Activations

Raw internal values, shown without further decomposition.

Field term · Standard throughout the interpretability literature; no single coining paper

Architecture

The shape of the model itself: what parts it has, and the order they run in.

The model's structure rather than anything it learned.

Common ML usage · General machine-learning usage; no single source

Attention

How the model decides which earlier words matter when it is working on the current one.

Which positions a model attends to, and how strongly.

Field term · Vaswani, A. et al., 2017. Attention Is All You Need

attention-heads

The parts that decide which earlier words matter for the next one.

What an attention head does, on a chosen task or across the full data distribution.

Field term · Elhage et al., 2021, A Mathematical Framework for Transformer Circuits, Anthropic.

attention-visualisation

Showing which words the model looked at, and how hard.

Displaying which positions a model attends to, and how strongly.

Common ML usage · Descriptive; earliest widely used tool is BertViz, Vig, 2019

attribution

Working out which inputs were responsible for an output.

Assigning responsibility for an output to particular inputs or components.

Common ML usage · No single origin; saliency and attribution literature from Simonyan et al., 2014 and Sundararajan et al., 2017

Attribution

How much each bit of the input pushed the model towards its answer.

How much each part of an input pushed the model towards its output. Distinct from a feature, which is what the model represents, and from an activation, which is how hard a component is firing.

Field term · Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K. and Mordvintsev, A., 2018. The Building Blocks of Interpretability. Distill. Verified first-hand 2026-08-22

attribution-graph

A map of which internal pieces caused which, for one specific prompt.

A causal graph of interpretable components explaining a single forward pass.

Field term · Ameisen et al., 2025, 'Circuit Tracing', Anthropic

auto-interp

Using one model to write explanations of what is happening inside another.

Using a model to generate natural-language explanations of another model's internals.

Field term · Practice introduced by Bills et al., 2023, OpenAI; 'auto-interp' is the field's established shorthand.

bias

The model treating people or groups differently in ways it should not.

Systematic differential treatment of people or groups by the model.

Common ML usage · Widespread across ML fairness literature; no single origin

Both

The work does each in turn, typically proposing concepts to interpret a decomposition.

Field term · Sharkey et al., 2025 §2 — the paper's own framing

Both

The work observes to generate a hypothesis and intervenes to test it.

Field term · Kowalska & Kwasnicka, 2026 §2.2 — the paper's own taxonomy

causal-tracing

Finding where in the model a particular fact is stored.

Locating where within a network a specific piece of knowledge is stored.

Field term · Meng et al., 2022, 'Locating and Editing Factual Associations in GPT' (ROME) (as cited in Kowalska & Kwasnicka, 2026)

Circuits

A chain of parts inside the model that work together to do one job.

Connected components and the paths between them.

Field term · Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M. and Carter, S., 2020. Zoom In: An Introduction to Circuits. Distill

Concept-based

Proposing a human concept in advance and then locating where the network represents it.

Field term · Sharkey et al., 2025 §2 — the paper's own framing

concept-injection

Inserting a known idea into the model's internals to see whether it notices.

Inserting a known representation into an unrelated context to test detection or effect.

Field term · Lindsey, 2025, 'Emergent Introspective Awareness', Anthropic — names the technique

crosscoder

A variant that works across several layers, or several models, at once.

A sparse dictionary trained jointly across several layers or several models.

Field term · Lindsey et al., 2024, 'Sparse Crosscoders for Cross-Layer Features and Model Diffing', Anthropic (as cited in Sharkey et al., 2025)

Dashboard

A single screen putting several different readouts side by side.

Several visual forms arranged together as one interface, read as a unit.

Field term · Chen, Y. et al., 2024. Designing a Dashboard for Transparency and Control of Conversational AI. Also standard field usage: 'feature dashboard'

dimensionality-reduction

Flattening very high-dimensional data down to two or three dimensions so it can be drawn.

Projecting high-dimensional activations into two or three dimensions so their structure can be seen.

Common ML usage · Standard technique (PCA, t-SNE, UMAP); no interpretability-specific origin.

emotion

Internal states that work like emotions and change what the model does.

Representations of emotional states and their causal effect on behaviour.

Field term · 'Functional emotions' / 'emotion vectors': Sofroniew et al., 2026, Anthropic

evaluation-awareness

A model noticing it is being tested, and behaving differently because of it.

A model recognising that it is being tested, and behaving differently as a result.

Field term · In use across Anthropic system cards and Apollo Research, 2025. NO SINGLE COINING PAPER IDENTIFIED — verify before citing

feature-interpretation

Working out what one small piece of the model has learned to recognise.

What an individual learned feature means.

Field term · In use across Bricken et al., 2023 and Templeton et al., 2024; Sharkey et al., 2025 §2.1.3 frames it as 'describing the functional role of components'.

feature-visualisation

Generating an image that shows what a part of the model responds to most strongly.

Synthesising an input that maximally activates a chosen component.

Field term · Olah, Mordvintsev & Schubert, 2017, 'Feature Visualization', Distill

Features

A single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.

Individual learned directions or concepts inside the model.

Field term · Bricken, T. et al., 2023, Towards Monosemanticity; Olah, C. et al., 2020, Zoom In: An Introduction to Circuits, Distill

Flow diagram

Boxes joined by arrows, showing what happens in what order.

Boxes and arrows showing a process or structure in sequence.

Information visualisation · Deng, D., Wu, Y., Shu, X., Wu, J., Fu, S., Cui, W. and Wu, Y., 2022. VisImages: A Fine-Grained Expert-Annotated Visualization Dataset. IEEE TVCG. NOT READ FIRST-HAND — taxonomy identified via search summary, verify before publication

hallucination

The model stating something confidently when it has nothing to base it on.

Producing confident content for which the model has no basis.

Common ML usage · Widespread in NLP; no single coining source

Heatmap

A grid where stronger colour means a bigger number.

A grid in which colour intensity encodes magnitude.

Information visualisation · Deng, D., Wu, Y., Shu, X., Wu, J., Fu, S., Cui, W. and Wu, Y., 2022. VisImages: A Fine-Grained Expert-Annotated Visualization Dataset. IEEE TVCG. NOT READ FIRST-HAND — taxonomy identified via search summary, verify before publication

Highlighted text

Ordinary text with words shaded in, where the shading shows how strongly the model reacted to each word.

Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.

Field term · Anthropic, 2023-2024, Towards Monosemanticity and Scaling Monosemanticity: 'text examples where the feature activates most strongly, with activating tokens highlighted'. Also 'token highlighting' in the visual analytics literature

in-context-learning

The model picking up a new skill from the prompt alone, without being retrained.

Acquiring a capability from the prompt rather than from training.

Field term · Brown et al., 2020 (GPT-3); mechanism in Olsson et al., 2022

induction-heads

Parts of the model that spot a pattern earlier in the text and continue it.

Attention heads that continue a pattern seen earlier in the context.

Field term · Elhage et al., 2021, 'A Mathematical Framework for Transformer Circuits'; developed in Olsson et al., 2022, Anthropic

Interventional

The model is altered and the consequences observed, supporting causal claims.

Field term · Kowalska & Kwasnicka, 2026 §2.2 — the paper's own taxonomy

introspection

Whether a model can accurately report on its own thinking.

A model's capacity to report on its own internal states.

Field term · Lindsey, 2025, 'Emergent Introspective Awareness', Anthropic

jailbreak

Getting a model to do something it was trained to refuse.

Circumventing a model's trained refusals.

Common ML usage · Widespread; borrowed from software security usage

knowledge-editing

Deliberately changing one specific fact a model holds.

Deliberately altering a specific fact the model holds.

Field term · Meng et al., 2022 (ROME); survey Wang et al., 2024 (as cited in Sharkey et al., 2025 §3.2.2)

logit-lens

Peeking at what the model would say if it stopped thinking at this layer.

Projecting an intermediate representation directly into the output vocabulary.

Field term · nostalgebraist, 2020 (as cited in Sharkey et al., 2025 and Kowalska & Kwasnicka, 2026 §4.1.2)

Model behaviour

What the model does, observed from the outside.

Common ML usage · General machine-learning usage; no single source

Model differences

What separates one model from another.

Field term · Model diffing / crosscoder literature; no single coining paper identified

model-diffing

Comparing two models to isolate exactly what is different between them.

Comparing two models in order to isolate what differs between them.

Field term · Bricken et al., 2024, 'Stage-Wise Model Diffing'; Lindsey et al., 2024, Anthropic

Neither

The work does not attempt either, as with purely pedagogical artefacts.

Field term · Sharkey et al., 2025 §2 — the paper's own framing

Observational

The model is examined without alteration, so the findings are correlational.

Field term · Kowalska & Kwasnicka, 2026 §2.2 — the paper's own taxonomy

persona

The character a model plays, and how easily it slips out of it.

The character a model adopts, and the stability of that adoption.

Field term · 'Persona vectors': Anthropic, 2025, arXiv:2507.21509

planning

The model deciding where a sentence is going before it writes it.

Selecting a future output before generating the text that leads to it.

Common ML usage · General term; as an interpretability finding, Lindsey et al., 2025, 'On the Biology of a Large Language Model'

polysemanticity

One part of the model doing several unrelated jobs at once.

A single component responding to several unrelated concepts.

Field term · Olah et al., 2020, 'Zoom In', Distill; Sharkey et al., 2025 traces the observation to Olah et al., 2017 and earlier

probing

Training a small, simple classifier to test whether some information is present inside the model.

Training a simple classifier on internal activations to test what they encode.

Field term · Kohn, 2015; Gupta et al., 2015; Alain & Bengio, 2017 (as cited in Sharkey et al., 2025 §2.2.1)

refusal

How a model decides to say no.

The mechanism by which a model declines a request.

Field term · 'Refusal direction': Arditi et al., 2024, 'Refusal in Language Models Is Mediated by a Single Direction' (as cited in Sharkey et al., 2025)

Reverse-engineering

Decomposing the network into parts and then determining what each part does.

Field term · Sharkey et al., 2025 §2 — the paper's own framing

Scatter plot

A graph where each point represents a pair of values, showing the relationship between two sets of numbers on a coordinate plane.

A graph of paired numerical values, with one variable on the horizontal axis and the corresponding value of a second variable on the vertical axis, used to reveal relationships or association between the variables.

Information visualisation · NIST/SEMATECH, Engineering Statistics Handbook, section 1.3.3.26, Scatter Plot. https://itl.nist.gov/div898/handbook/eda/section3/eda33q.htm

sparse-autoencoder

A technique for pulling a model's tangled internals apart into separate, nameable pieces.

Learning an overcomplete, sparsely activating basis for a layer's activations.

Field term · Sparse coding: Olshausen & Field, 1997 (as cited in Kowalska & Kwasnicka, 2026). Applied to interpretability: Sharkey, Braun & Millidge, 2022; Bricken et al., 2023; Huben/Cunningham et al., 2024

Spatial attribution

Colour laid over a picture showing which parts of it the model was reacting to.

Attribution values rendered over the positions of an input image, showing which regions drove the output. A saliency map is the most common instance.

Field term · Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K. and Mordvintsev, A., 2018. The Building Blocks of Interpretability. Distill. Verified first-hand 2026-08-22

steering

Nudging the model's internals mid-thought to change what it says.

Adding or subtracting a direction in activation space in order to alter behaviour.

Field term · 'Activation steering' / 'activation addition': Turner et al., 2024, credited with introducing it in Sharkey et al., 2025 §3.2.2

superposition

A model packing more ideas into its wiring than it has room for, by letting them overlap.

Representing more features than there are dimensions, by encoding them sparsely and near-orthogonally.

Field term · Elhage et al., 2022, 'Toy Models of Superposition', Anthropic (as cited in Sharkey et al., 2025)

Training dynamics

How the model changes over the course of training.

Common ML usage · General machine-learning usage; no single source

transcoder

A variant that reconstructs what a layer passes onward, rather than what it received.

A sparse dictionary trained to reproduce a layer's output rather than its input.

Field term · Dunefsky, Chlenski & Nanda, 2024 (as cited in Sharkey et al., 2025)

unfaithful-reasoning

The explanation a model gives not matching the reasoning it actually did.

Stated reasoning that does not correspond to the computation actually performed.

Field term · 'Chain-of-thought unfaithfulness': Turpin et al., 2023; Arcuschin, Conmy et al., 2025 (as cited in Sharkey et al., 2025)

User model

The picture a model builds of who it is talking to, such as age, gender or mood.

What the model has inferred about the person using it.

Field term · Chen, Y. et al., 2024. Designing a Dashboard for Transparency and Control of Conversational AI

user-modelling

What the model has quietly worked out about the person talking to it.

The model's internal inferences about the person it is talking to.

Field term · 'User model': Chen et al., 2024, TalkTuner, arXiv:2406.07882

world-models

Signs that a model has built an internal picture of something nobody told it directly.

Internal representations of an external state the model was never given directly.