Academia
A university or academic research group.
A university or academic research group.
The level of model access a person would need to reproduce the work.
Requires gated or paid inference access, but not weights.
Preserved in static or historical form and no longer maintained.
Work made primarily for aesthetic or critical reception rather than scientific reporting.
The form the work takes, independent of what it contains or argues.
A practitioner working primarily within an art context.
Coined for this index to describe its own structure. Not a claim about the field, and no external source exists.
Models that process sound or speech.
Supplied by Ava in conversation rather than drawn from a source.
Work seeking to reverse-engineer the model's general structure, irrespective of any single behaviour.
Work pursuing a specific behaviour in order to establish general structure, or the reverse.
A searchable collection of results presented for inspection.
The named individual contributors, in the order the source credits them.
A full bibliographic reference in a consistent style.
The practice is real but the LABEL was invented by Claude and is not established. Rename before publishing where a standard term exists.
Models whose subject matter is programming languages.
A link to that source code.
In wide use across machine learning with no single coining source. Safe to use, not attributable.
How well evidenced this row is, judged by what the person writing it actually read.
A value that may only be chosen from the list on the Vocabularies sheet, never typed freely.
Where the producing institution is based.
A live readout of a system's internal state, displayed alongside that system in use.
When the row was created and by whom, human or model.
The date the artefact first became publicly available, in ISO form where known.
The URL no longer resolves to the artefact.
A model made publicly usable in a deliberately modified state, as a demonstration.
One or two sentences describing the artefact to a stranger, written to be published rather than as a note to self.
The visual exists so that a reader can inspect and audit a claim which is itself stated in prose.
The artefact is undergone rather than read, and its value lies in what it does to the visitor.
The visual communicates a result established elsewhere, serving teaching rather than discovery.
Established vocabulary in the published interpretability literature. The citation is the earliest use Claude could verify, not a guaranteed first use.
The result is constituted by the visual, such that removing the images would remove the contribution.
A field written in prose, with no fixed list of permitted values.
An organisation training state-of-the-art general-purpose models.
Joint authorship spanning both.
Not tied to one kind of data. Applies across different models and media.
Not tied to any particular model or data type.
Non-specialists, with no assumed technical background.
Models of genetic sequence.
The primary source was read in full by whoever wrote the row.
Origin says who coined the TERM. Coined in gives the citation. Definition drafted from says what Claude actually read to write the sentence — which is not the same thing.
A permanent identifier for one artefact, never reused, so that a retired entry's number can never resolve to a different record.
Models that produce images, such as diffusion models.
A person working outside any institution.
Others can operate it on their own inputs or models to produce findings that did not previously exist.
Who the artefact was designed to be understood by.
What a person can actually do with the artefact, verb by verb.
A written argument in which the figures respond to the reader.
A standalone application whose purpose is to let a user investigate a model.
Requires access available only within the producing organisation.
Text. Models that read and write words.
Models that generate or process text.
Whether the recorded URL currently resolves to the artefact.
The URL resolves and the artefact functions as described.
The row rests on inference and has not been checked against a source.
The thing rendered on screen. Not what the work is about, which is the phenomenon tag, but what you are actually looking at.
The image filenames belonging to this entry, semicolon separated.
The row rests on an abstract, documentation, or a reliable secondary source.
The interpretability technique the work employs.
The interpretability technique or techniques the artefact rests on.
Deliberately designed to serve more than one of the above.
The kind of model or data the work concerns.
The specific neural networks the work examines.
Models that process more than one modality jointly.
The artefact's own title, as given by its makers.
Work that neither diagnoses a behaviour nor uncovers structure, but conveys existing knowledge.
No source is published.
The artefact was operable once but is not now.
An organisation constituted for public benefit rather than profit.
This entry produces no picture of its own.
No graphic form. The artefact qualifies some other way.
Requires no model access at all; the outputs are fixed.
Code published to be run and adapted by other researchers.
Anything qualifying the row, including unresolved questions and flagged boundary cases.
There is nothing visual to look at.
Nothing is rendered. The artefact is encountered rather than viewed.
Whether the work examines the model unchanged or alters it and observes the consequence.
A visitor can supply their own prompt, image or model and see fresh output.
Whether the artefact's own source code is publicly available.
Reproducible using publicly downloadable model weights.
The institution that produced the work.
A form not covered above; propose a new term rather than reusing this indefinitely.
Some components are released and others deliberately withheld.
Reachable only behind payment or registration.
The property or behaviour of the model that the work is about.
Infrastructure hosting many artefacts, models or datasets originating from multiple parties.
Regulators, governments and civil society organisations.
People deploying or auditing models professionally.
Work targeting a specific model behaviour, asking which parts of the network produce it.
Whether the work targets one specific behaviour or the model's general structure.
The visuals are fixed; interaction is limited to navigation such as panning or zooming.
A visitor can explore interactively, but only within cases the authors selected in advance.
Models of protein sequence or structure.
Every column in the Ledger and every permitted value is defined here in one sentence, so that two people filling in the same row reach the same answer.
The IDs of other Atlas entries with a documented relationship to this one.
Specialists in interpretability or machine learning.
Whether the work decomposes the network and then names the parts, or proposes concepts and then locates them.
Agents trained by reinforcement learning to act within an environment.
The function the visual element performs within the research, which is the central judgement in each row.
Whether, and how far, a person arriving today can operate the artefact themselves.
Parts of the model to do with harm, deception, bias or misuse.
Representations bearing on harm, deception, bias or misuse.
The kind of institution that produced the work, recorded so the composition of the field is visible.
The lowercase hyphenated form of the entry's name, used as its web address.
What was read to write this row, and on what date.
A commercial company whose product is interpretability or model tooling.
Fixed images within a conventional publication.
Learners being taught the material for the first time.
Models over structured, columnar data.
Faceted keywords in facet:value form, drawn only from the controlled facet lists.
An image representing the entry.
The role has not been determined.
Not determined, often because the page could not be loaded.
The basis for the row was not recorded.
Not yet determined.
Not yet determined.
The canonical public address of the artefact itself, rather than of a paper describing it.
Where the work appeared, which is often not the same as who produced it.
Time-based linear media.
Images. Models that look at pictures.
Models that process static images.
Observations on the artefact's visual and interaction design, including what makes it work or fail.
The shape the picture takes: what you would call the graphic if describing it to someone who could not see it.
The object being made visible, stated in plain terms a non-specialist could follow.
One line stating the artefact's significance in terms a reader could dispute.
A tag is a filter, not an opinion; claims about significance belong in Why it matters, where a reader can dispute them.
Organisation, artefact type, sector and open source are already columns; tagging them again would create two sources of truth that drift apart.
The publication year held separately, so that entries sort chronologically without parsing dates.
The artefact's own source code is publicly available under a licence.
Lots of tiny pictures in a grid, each showing what the model saw at that spot.
Many small images laid out in a grid, each one standing for what a part of the model responded to at that position.
Swapping one internal value for another to see whether it mattered.
Substituting one activation for another to measure its causal effect on the output.
Raw internal values, shown without further decomposition.
The shape of the model itself: what parts it has, and the order they run in.
The model's structure rather than anything it learned.
How the model decides which earlier words matter when it is working on the current one.
Which positions a model attends to, and how strongly.
The parts that decide which earlier words matter for the next one.
What an attention head does, on a chosen task or across the full data distribution.
Showing which words the model looked at, and how hard.
Displaying which positions a model attends to, and how strongly.
Working out which inputs were responsible for an output.
Assigning responsibility for an output to particular inputs or components.
How much each bit of the input pushed the model towards its answer.
How much each part of an input pushed the model towards its output. Distinct from a feature, which is what the model represents, and from an activation, which is how hard a component is firing.
A map of which internal pieces caused which, for one specific prompt.
A causal graph of interpretable components explaining a single forward pass.
Using one model to write explanations of what is happening inside another.
Using a model to generate natural-language explanations of another model's internals.
The model treating people or groups differently in ways it should not.
Systematic differential treatment of people or groups by the model.
The work does each in turn, typically proposing concepts to interpret a decomposition.
The work observes to generate a hypothesis and intervenes to test it.
Finding where in the model a particular fact is stored.
Locating where within a network a specific piece of knowledge is stored.
A chain of parts inside the model that work together to do one job.
Connected components and the paths between them.
Proposing a human concept in advance and then locating where the network represents it.
Inserting a known idea into the model's internals to see whether it notices.
Inserting a known representation into an unrelated context to test detection or effect.
A variant that works across several layers, or several models, at once.
A sparse dictionary trained jointly across several layers or several models.
A single screen putting several different readouts side by side.
Several visual forms arranged together as one interface, read as a unit.
Flattening very high-dimensional data down to two or three dimensions so it can be drawn.
Projecting high-dimensional activations into two or three dimensions so their structure can be seen.
Internal states that work like emotions and change what the model does.
Representations of emotional states and their causal effect on behaviour.
A model noticing it is being tested, and behaving differently because of it.
A model recognising that it is being tested, and behaving differently as a result.
Working out what one small piece of the model has learned to recognise.
What an individual learned feature means.
Generating an image that shows what a part of the model responds to most strongly.
Synthesising an input that maximally activates a chosen component.
A single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.
Individual learned directions or concepts inside the model.
Boxes joined by arrows, showing what happens in what order.
Boxes and arrows showing a process or structure in sequence.
The model stating something confidently when it has nothing to base it on.
Producing confident content for which the model has no basis.
A grid where stronger colour means a bigger number.
A grid in which colour intensity encodes magnitude.
Ordinary text with words shaded in, where the shading shows how strongly the model reacted to each word.
Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.
The model picking up a new skill from the prompt alone, without being retrained.
Acquiring a capability from the prompt rather than from training.
Parts of the model that spot a pattern earlier in the text and continue it.
Attention heads that continue a pattern seen earlier in the context.
The model is altered and the consequences observed, supporting causal claims.
Whether a model can accurately report on its own thinking.
A model's capacity to report on its own internal states.
Getting a model to do something it was trained to refuse.
Circumventing a model's trained refusals.
Deliberately changing one specific fact a model holds.
Deliberately altering a specific fact the model holds.
Peeking at what the model would say if it stopped thinking at this layer.
Projecting an intermediate representation directly into the output vocabulary.
What the model does, observed from the outside.
What separates one model from another.
Comparing two models to isolate exactly what is different between them.
Comparing two models in order to isolate what differs between them.
The work does not attempt either, as with purely pedagogical artefacts.
Dots joined by lines, showing what connects to what.
Nodes joined by edges, showing what connects to or causes what.
The model is examined without alteration, so the findings are correlational.
The character a model plays, and how easily it slips out of it.
The character a model adopts, and the stability of that adoption.
The model deciding where a sentence is going before it writes it.
Selecting a future output before generating the text that leads to it.
One part of the model doing several unrelated jobs at once.
A single component responding to several unrelated concepts.
Training a small, simple classifier to test whether some information is present inside the model.
Training a simple classifier on internal activations to test what they encode.
How a model decides to say no.
The mechanism by which a model declines a request.
Decomposing the network into parts and then determining what each part does.
A graph where each point represents a pair of values, showing the relationship between two sets of numbers on a coordinate plane.
A graph of paired numerical values, with one variable on the horizontal axis and the corresponding value of a second variable on the vertical axis, used to reveal relationships or association between the variables.
A technique for pulling a model's tangled internals apart into separate, nameable pieces.
Learning an overcomplete, sparsely activating basis for a layer's activations.
Colour laid over a picture showing which parts of it the model was reacting to.
Attribution values rendered over the positions of an input image, showing which regions drove the output. A saliency map is the most common instance.
Nudging the model's internals mid-thought to change what it says.
Adding or subtracting a direction in activation space in order to alter behaviour.
A model packing more ideas into its wiring than it has room for, by letting them overlap.
Representing more features than there are dimensions, by encoding them sparsely and near-orthogonally.
How the model changes over the course of training.
A variant that reconstructs what a layer passes onward, rather than what it received.
A sparse dictionary trained to reproduce a layer's output rather than its input.
The explanation a model gives not matching the reasoning it actually did.
Stated reasoning that does not correspond to the computation actually performed.
The picture a model builds of who it is talking to, such as age, gender or mood.
What the model has inferred about the person using it.
What the model has quietly worked out about the person talking to it.
The model's internal inferences about the person it is talking to.
Signs that a model has built an internal picture of something nobody told it directly.
Internal representations of an external state the model was never given directly.