MATH · IN · MODELS

Methods (242)

Grouped by what kind of methodological move they make — not by which paper happened to name them. Each card is a category; open it to see the techniques inside and where they've been used.

Causal Validation

A direction or feature that merely correlates with a concept isn't enough — these methods intervene on the model's internals and check whether behavior changes as predicted, establishing that the model actually uses the structure.
47methods · 252 papers →

Theoretical / Analytical

Reason about representation geometry mathematically — proofs, derivations, formal relationships between quantities — rather than collecting new empirical measurements.
43methods · 170 papers →

Direction Extraction

Turn a labeled contrast between two classes of activations into a single candidate feature direction — the raw material Linear Direction and the Linear Representation Hypothesis are tested against.
36methods · 264 papers →

Dictionary Learning

Unsupervised, sparse decomposition of activations into an overcomplete basis of candidate features — used when you don't already know what to probe for.
13methods · 82 papers →

Representation Alignment

Compare two representations' overall geometry to each other directly — do two models, modalities, or checkpoints measure similarity between the same datapoints in the same way — rather than analyzing the geometry of a single representation in isolation.
13methods · 40 papers →

Dimensionality Reduction

Project high-dimensional activations onto a low-dimensional basis to make geometric structure (manifolds, clusters, trajectories) visible and measurable.
12methods · 123 papers →

Corpus Statistics

Analyze the training corpus's word co-occurrence statistics directly — not model activations — to explain why a representational geometry arises in the first place.
2methods · 6 papers →