Methods (242)
Grouped by what kind of methodological move they make — not by which paper happened to name them. Each card is a category; open it to see the techniques inside and where they've been used.
Causal Validation
A direction or feature that merely correlates with a concept isn't enough — these methods intervene on the model's internals and check whether behavior changes as predicted, establishing that the model actually uses the structure.
47methods · 252 papers →
Theoretical / Analytical
Reason about representation geometry mathematically — proofs, derivations, formal relationships between quantities — rather than collecting new empirical measurements.
43methods · 170 papers →
Direction Extraction
Turn a labeled contrast between two classes of activations into a single candidate feature direction — the raw material Linear Direction and the Linear Representation Hypothesis are tested against.
36methods · 264 papers →
Dictionary Learning
Unsupervised, sparse decomposition of activations into an overcomplete basis of candidate features — used when you don't already know what to probe for.
13methods · 82 papers →
Representation Alignment
Compare two representations' overall geometry to each other directly — do two models, modalities, or checkpoints measure similarity between the same datapoints in the same way — rather than analyzing the geometry of a single representation in isolation.
13methods · 40 papers →
Dimensionality Reduction
Project high-dimensional activations onto a low-dimensional basis to make geometric structure (manifolds, clusters, trajectories) visible and measurable.
12methods · 123 papers →
Corpus Statistics
Analyze the training corpus's word co-occurrence statistics directly — not model activations — to explain why a representational geometry arises in the first place.
2methods · 6 papers →