method
sparse-autoencoder
4 entries
2024 (2)
IA-008Golden Gate Claude has no visual form and shows nothingThere is nothing visual to look at.Nothing is rendered. The artefact is encountered rather than viewed.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
IA-003Scaling Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and safety-relevant-featuresParts of the model to do with harm, deception, bias or misuse.Representations bearing on harm, deception, bias or misuse.glossary.
2023 (2)
IA-002Towards Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about superpositionA model packing more ideas into its wiring than it has room for, by letting them overlap.Representing more features than there are dimensions, by encoding them sparsely and near-orthogonally.glossary, polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
IA-006
Neuronpedia is a dashboardA single screen putting several different readouts side by side.Several visual forms arranged together as one interface, read as a unit.glossary and a node-link graphDots joined by lines, showing what connects to what.Nodes joined by edges, showing what connects to or causes what.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary and circuitsA chain of parts inside the model that work together to do one job.Connected components and the paths between them.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
