MATH · IN · MODELS
methods / Causal Validation / Activation patching

Activation patching

Techniqueintermediate

Replaces (patches) an activation from one forward pass into another to causally test which components carry a given piece of information — an interchange intervention, not simple vector arithmetic.

Used in (40 observations)

structure: 1D continuum manifold · models: Gemma-2-2B, Gemma-2-9B, Qwen3-4B · paper: Geometry of Ordinal Representations in Language Models, When Models Manipulate Manifolds: The Geometry of a Counting Task
structure: Linear Separability · models: Qwen3-32B · paper: Tool-Call Dependency Structure Is Linearly Decodable in LLM Agent Residual Streams
structure: Linear Subspace · models: Llama-3-8B-Instruct, Qwen3-8B · paper: Cell-Based Representation of Relational Binding in Language Models
structure: Linear Direction · models: Qwen3-VL-4B-Instruct, Qwen3-VL-8B-Instruct, LLaVA-OneVision-1.5-4B-Instruct · paper: Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
structure: Linear Direction, Linear Subspace · models: Gemma 3 4B Instruct, Gemma 3 4B, Gemma 3 27B Instruct, Qwen3-4B-Instruct, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Tool Calling Is Linearly Readable and Steerable in Language Models
structure: Tree Metric Embedding · models: Gemma 3 27B, Llama-3.1-8B, Mistral-7B-v0.3, Qwen3-14B, Qwen3-4B · paper: A Minimalist Structural Probe for Phase-Count Abstraction in Transformer Language Models
structure: Cone · models: Llama-3.1-8B-Instruct, Llama-3-8B-Instruct, Llama-3.2-3B-Instruct, Gemma-2-2B-it, Gemma-2-9B-it, OPT-1.3B, OPT-2.7B, OPT-6.7B, Pythia-410M, Pythia-1.4B, Pythia-2.8B, Mistral-7B-Instruct-v0.1 · paper: Investigating Representation Universality: Case Study on Genealogical Representations
structure: Linear Direction · models: Qwen2-VL-7B-Instruct, Gemma 3 4B Instruct · paper: The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
structure: Linear Direction · models: LLaMA-7B · paper: Label Words as Local Task Vectors in In-Context Learning
structure: Linear Direction, Linear Subspace · models: LLaMA-30B, LLaMA-13B, LLaMA-65B, Pythia-2.8B · paper: How Do Language Models Bind Entities in Context?
structure: Generalized Helix, Circle · models: Llama-3.1-8B · paper: Arithmetic in the Wild: Llama Uses Base-10 Addition to Reason About Cyclic Concepts
structure: Linear Direction · models: XGLM-7.5B, EuroLLM-9B, mT5-xl · paper: How Do Multilingual Language Models Remember Facts?
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B, Gemma-2-27B, Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B-Instruct · paper: How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
structure: Linear Direction · models: GPT-J-6B, GPT-NeoX-20B, Llama-2-7B, Llama-2-13B, Llama-2-70B, LLaMA-7B, LLaMA-13B, LLaMA-30B, Pythia-2.8B, Pythia-6.9B, Pythia-12B · paper: Function Vectors in Large Language Models, In-Context Learning Creates Task Vectors, Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning, Do Different Prompting Methods Yield a Common Task Representation in Language Models?
structure: Circle, Linear Direction · models: V-JEPA 2, VideoMAE-v2 · paper: Interpreting Physics in Video World Models
structure: Linear Direction · models: LLaVA-1.5-7B, Llama-3.1-8B-Instruct, InternVL3-8B, Gemma-7B-it · paper: Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
structure: Generalized Helix · models: GPT-J-6B, Pythia-6.9B, Llama-3.1-8B · paper: Language Models Use Trigonometry to Do Addition
structure: Linear Direction · models: EuroLLM-1.7B · paper: When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training
structure: Linear Direction, Linear Separability · models: Gaperon-8B · paper: Language-Switching Triggers Take a Latent Detour Through Language Models
structure: Linear Direction · models: AudioLDM2, Stable Audio Open, Ace-Step · paper: TADA! Tuning Audio Diffusion Models Through Activation Steering
structure: Linear Subspace · models: Custom Conv1D-encoder + 3-layer BiGRU + HiFi-GAN brain-to-speech decoder (trained on real human sEEG, VOCALMIND) · paper: Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Gemma-2-9B, Gemma-2-9B-it, Mistral-7B-v0.3, Mistral-7B-Instruct-v0.3 · paper: Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens
structure: Linear Direction · models: Gemma 2 9B IT, Gemma 3 12B IT, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct · paper: A Retrieval-Conditioned Rebinding Circuit for Dynamic Entity Tracking in Large Language Models
structure: Linear Direction, Polytope (Simplex) · models: Pythia-160M · paper: (How) Do Language Models Track State?
structure: Linear Subspace · models: PhonSSM (AGAN + PDM + bidirectional Mamba SSM + HPC, skeleton-based ASL) · paper: State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition
structure: Linear Subspace · models: Qwen2.5-14B-Instruct, Llama-3-70B-Instruct, Llama-3.1-405B-Instruct · paper: Language Models Use Lookbacks to Track Beliefs
structure: Relation frame (ordered multi-token tuple geometry) · models: Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Llama-3.1-405B-Instruct · paper: Relational Rank Geometry in Transformers: Detecting and Steering Hidden-State Relation Frames
structure: 1D continuum manifold · models: Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Claude 3.5 Haiku · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, When Models Manipulate Manifolds: The Geometry of a Counting Task
structure: Linear Direction · models: ViT-B/16 (ImageNet-21k, supervised) · paper: From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers
structure: Linear Direction · models: GPT-2 Small, Pythia-1.4B, Pythia-2.8B · paper: Linear Representations of Sentiment in Large Language Models
structure: Linear Direction · models: GPT-2 Small, Pythia-2.8B · paper: Linear Representations of Sentiment in Large Language Models
structure: Linear Subspace, Linear Direction · models: Llama-3-8B, Llama-3.1-8B, Llama-3.2-3B, Qwen2-7B (base), Qwen2.5-32B, Yi-34B (base) · paper: Task Recognition and Task Learning Heads Align In-Context Hidden States with a Label-Unembedding Task Subspace
structure: Linear Direction · models: Qwen3-4B-Instruct-2507 · paper: Temporal Preference Concepts and Their Functions in a Large Language Model
structure: Circle, Cone, Platonic Representation Hypothesis · models: GPT-2-small, Mistral-7B, Llama-3-8B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Llama-3.1-8B · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, Not All Language Model Features Are One-Dimensionally Linear, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Do Sparse Autoencoders Capture Concept Manifolds?
structure: Linear Direction · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-v0.1 · paper: The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets, On the Universal Truthfulness Hyperplane Inside LLMs
structure: Linear Direction · models: Llama-3.2-1B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct · paper: Negative Before Positive: Asymmetric Valence Processing in Large Language Models
structure: Linear Subspace · models: Qwen2.5-VL 7B Instruct, InternVL3-8B · paper: Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
structure: Linear Direction, Constraint-Algebra Basis Hypothesis · models: Sudoku Transformer (8 layers, 8 heads, d_model=576) · paper: Transformers Linearly Represent Highly Structured World Models
structure: Linear Subspace · models: 12-layer, 8-head causal Transformer (d_model=512, RoPE, synthetic variable-assignment-program task) · paper: How Do Transformers Learn Variable Binding in Symbolic Programs?