method
steering
4 entries
2024 (3)
IA-011TalkTuner is a dashboardA single screen putting several different readouts side by side.Several visual forms arranged together as one interface, read as a unit.glossary showing a user modelThe picture a model builds of who it is talking to, such as age, gender or mood.What the model has inferred about the person using it.glossary, about biasThe model treating people or groups differently in ways it should not.Systematic differential treatment of people or groups by the model.glossary and user-modellingWhat the model has quietly worked out about the person talking to it.The model's internal inferences about the person it is talking to.glossary.
IA-008Golden Gate Claude has no visual form and shows nothingThere is nothing visual to look at.Nothing is rendered. The artefact is encountered rather than viewed.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
IA-003Scaling Monosemanticity is highlighted textOrdinary text with words shaded in, where the shading shows how strongly the model reacted to each word.Running text with individual words or tokens shaded to encode a per-token value, most often how strongly a feature activated there.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary, about polysemanticityOne part of the model doing several unrelated jobs at once.A single component responding to several unrelated concepts.glossary and safety-relevant-featuresParts of the model to do with harm, deception, bias or misuse.Representations bearing on harm, deception, bias or misuse.glossary.
2023 (1)
IA-006
Neuronpedia is a dashboardA single screen putting several different readouts side by side.Several visual forms arranged together as one interface, read as a unit.glossary and a node-link graphDots joined by lines, showing what connects to what.Nodes joined by edges, showing what connects to or causes what.glossary showing featuresA single thing the model has learned to recognise, like a place, a tone of voice, or a kind of mistake.Individual learned directions or concepts inside the model.glossary and circuitsA chain of parts inside the model that work together to do one job.Connected components and the paths between them.glossary, about feature-interpretationWorking out what one small piece of the model has learned to recognise.What an individual learned feature means.glossary.
