artefact type
Interactive article
5 entries
2025 (1)
IA-004On the Biology of a Large Language Model is a node-link graphDots joined by lines, showing what connects to what.Nodes joined by edges, showing what connects to or causes what.glossary showing circuitsA chain of parts inside the model that work together to do one job.Connected components and the paths between them.glossary, about planningThe model deciding where a sentence is going before it writes it.Selecting a future output before generating the text that leads to it.glossary, hallucinationThe model stating something confidently when it has nothing to base it on.Producing confident content for which the model has no basis.glossary, jailbreakGetting a model to do something it was trained to refuse.Circumventing a model's trained refusals.glossary and unfaithful-reasoningThe explanation a model gives not matching the reasoning it actually did.Stated reasoning that does not correspond to the computation actually performed.glossary.
2024 (1)
IA-003Scaling Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and safety-relevant-featuresParts of the model to do with harm, deception, bias or misuse.Representations bearing on harm, deception, bias or misuse.glossary.
2023 (1)
IA-002Towards Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about superpositionA model packing more ideas into its wiring than it has room for, by letting them overlap.Representing more features than there are dimensions, by encoding them sparsely and near-orthogonally.glossary, polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
2019 (1)
IA-007
Activation Atlas is a scatter plotA graph where each point represents a pair of values, showing the relationship between two sets of numbers on a coordinate plane.A graph of paired numerical values, with one variable on the horizontal axis and the corresponding value of a second variable on the vertical axis, used to reveal relationships or association between the variables.glossary and an activation gridLots of tiny pictures in a grid, each showing what the model saw at that spot.Many small images laid out in a grid, each one standing for what a part of the model responded to at that position.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
