makes visible
Features
5 entries
2024 (1)
IA-003Scaling Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and safety-relevant-featuresParts of the model to do with harm, deception, bias or misuse.Representations bearing on harm, deception, bias or misuse.glossary.
2023 (2)
IA-002Towards Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about superpositionA model packing more ideas into its wiring than it has room for, by letting them overlap.Representing more features than there are dimensions, by encoding them sparsely and near-orthogonally.glossary, polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
IA-006
Neuronpedia is a dashboardA single screen putting several different readouts side by side.Several visual forms arranged together as one interface, read as a unit.glossary and a node-link graphDots joined by lines, showing what connects to what.Nodes joined by edges, showing what connects to or causes what.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary and circuitsA chain of parts inside the model that work together to do one job.Connected components and the paths between them.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.

2019 (1)
IA-007
Activation Atlas is a scatter plotA graph where each point represents a pair of values, showing the relationship between two sets of numbers on a coordinate plane.A graph of paired numerical values, with one variable on the horizontal axis and the corresponding value of a second variable on the vertical axis, used to reveal relationships or association between the variables.glossary and an activation gridLots of tiny pictures in a grid, each showing what the model saw at that spot.Many small images laid out in a grid, each one standing for what a part of the model responded to at that position.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
