MATH · IN · MODELS
methods / Causal Validation / Causal interventions (steering) / Activation Steering (Addition)

Activation Steering (Addition)

Techniqueintermediate

The additive half of causal-interventions, pulled out as its own citable technique: add a fixed direction to activations at inference time, x ↦ x + αr, to induce or strengthen a concept/behavior — as distinct from ablating it out.

Used in (74 observations)

structure: Linear Subspace · models: Llama-3.2-3B, Llama 3.1 70B, Qwen3-1.7B, Qwen3-32B · paper: Semantic Structure of Feature Space in Large Language Models
structure: Linear Direction · models: Qwen3-8B, GPT-OSS-20B, Qwen2.5-7B-Instruct · paper: What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
structure: Linear Subspace · models: Llama-3.1-8B-Instruct, Qwen2-7B-Instruct · paper: Death by a Thousand Directions: Exploring the Geometry of Harmfulness in LLMs through Subconcept Probing
structure: Linear Subspace, Linear Direction · models: Gemma 2 27B Instruct, Gemma-2-27B, Qwen3-32B, Llama 3.3 70B Instruct · paper: The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
structure: Linear Subspace, Linear Direction · models: Qwen3-8B, Llama-3.1-8B-Instruct · paper: The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models
structure: Linear Direction · models: Bielik-11B, Llama-PLLuM-12B, Gemma 4 12B, Mistral NeMo 12B (base), Llama-3.1-8B, Qwen3-14B · paper: Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering
structure: Linear Direction · models: Qwen3-VL-4B-Instruct, Qwen3-VL-8B-Instruct, LLaVA-OneVision-1.5-4B-Instruct · paper: Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
structure: Linear Direction, Linear Subspace · models: Gemma 3 4B Instruct, Gemma 3 4B, Gemma 3 27B Instruct, Qwen3-4B-Instruct, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Tool Calling Is Linearly Readable and Steerable in Language Models
structure: Circle, 1D continuum manifold · models: Llama-3.1-8B · paper: Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
structure: 1D continuum manifold, Linear Direction · models: Gemma-2-9B, Mistral-7B, Llama-3-70B-Instruct · paper: Latent Structure of Affective Representations in Large Language Models
structure: Linear Direction · models: VideoMAE-base · paper: Causal Physics Steering in Video World Models via Concept Activation Vectors
structure: Linear Direction · models: Llama-3.2-3B-Instruct, Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, Gemma 3 4B Instruct, Qwen2.5-7B-Instruct · paper: Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
structure: Linear Direction · models: Whisper small, Whisper large-v3 · paper: Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
structure: Linear Direction · models: DDPM++ (CelebA-HQ 256x256), DDPM++ (LSUN-Church 256x256), DDPM++ (LSUN-Bedroom 256x256), iDDPM (AFHQ-Dog 256x256), ADM P2-weighted (METFACES 256x256) · paper: Diffusion Models already have a Semantic Latent Space
structure: Linear Direction · models: Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, Vicuna-7B-v1.5 · paper: Linear Representations of Political Perspective Emerge in Large Language Models
structure: Linear Direction · models: Gemma 3 12B Instruct, Llama-3.1-8B-Instruct · paper: Dissociating the Internal Representations of Sycophancy in LLMs
structure: Conceptual Belief Space Hypothesis, Linear Subspace · models: Llama-3.1-8B-Instruct · paper: Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
structure: Linear Direction · models: Llama-2-7B-Chat · paper: Understanding (Un)Reliability of Steering Vectors in Language Models
structure: Linear Direction · models: Stable Diffusion v1.5 · paper: SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
structure: Polytope (Simplex) · models: Gemma-2B, Llama-3-8B, Qwen3-4B, Mistral-7B-v0.3 · paper: The Geometry of Categorical and Hierarchical Concepts in Large Language Models, When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs
structure: Linear Direction · models: Custom Transformer-VAE, Autoregressive MultiSlotting variant (trained from scratch on SELFIES-tokenized molecules) · paper: Molecules Meet Language: Confound-Aware Representation Learning and Chemical Property Steering in Transformer-VAE Latent Spaces
structure: Linear Direction · models: Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Persona Vectors: Monitoring and Controlling Character Traits in Language Models
structure: Anisotropy · models: Wan2.1-1.3B, Wan2.2-5B, CogVideoX-5B · paper: Steering Video Diffusion Transformers with Massive Activations
structure: Linear Direction, Linear Separability · models: Llama-3.2-1B, Llama-3.2-3B, Gemma-2-2B, Qwen2.5-1.5B, Llama-3.1-8B, Mistral-7B-v0.3 · paper: Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations
structure: Linear Direction · models: ChessGPT-8L-25M, ChessGPT-16L-50M · paper: Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
structure: Linear Direction · models: Claude 3 Sonnet · paper: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
structure: Linear Direction · models: Claude Sonnet 4.5 · paper: Emotion Concepts and their Function in a Large Language Model
structure: Paraboloid (circular × pinched continuum) · models: Llama-3.1-8B · paper: Do Sparse Autoencoders Capture Concept Manifolds?
structure: Conceptor (soft ellipsoidal region) · models: GPT-J-6B, GPT-NeoX-20B · paper: Steering Large Language Models using Conceptors: Improving Addition-Based Activation Engineering
structure: Linear Subspace, Linear Direction · models: GPT-2 Small, GPT-2 Medium, GPT-2 Large, Qwen2-1.5B-Instruct, Qwen2-7B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Gemma-2-2B-it · paper: The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
structure: Affine Subspace · models: Llama-3.1-8B · paper: Do Sparse Autoencoders Capture Concept Manifolds?
structure: Linear Direction · models: Qwen2-VL-7B-Instruct, Gemma 3 4B Instruct · paper: The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
structure: Linear Direction · models: Qwen2.5-VL 7B Instruct, Qwen3-VL 8B, Qwen3-VL 32B, LLaVA-OneVision-1.5 8B, Gemma 3 4B Instruct, Gemma 3 27B Instruct · paper: Causal Probing for Internal Visual Representations in Multimodal Large Language Models
structure: Linear Direction · models: DDPM (CelebA-HQ 256x256, HuggingFace Diffusers), DDPM+P2-weighting (AFHQ 256x256), DDPM (LSUN-Church 256x256, HuggingFace Diffusers), DDPM (LSUN-Bedroom 256x256, HuggingFace Diffusers), DDPM (LSUN-Cat 256x256), DDPM (LSUN-Horse 256x256), DDPM (ImageNet 256x256), DDPM+P2-weighting (FFHQ 256x256), DDPM+P2-weighting (Flowers 256x256), Stable Diffusion v2.1 · paper: Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry
structure: Generalized Helix, Circle · models: Llama-3.1-8B · paper: Arithmetic in the Wild: Llama Uses Base-10 Addition to Reason About Cyclic Concepts
structure: Linear Direction · models: Llama 3.3 70B Instruct · paper: Linear Personality Probing and Steering in LLMs: A Big Five Study
structure: Linear Direction · models: GPT-J-6B, GPT-NeoX-20B, Llama-2-7B, Llama-2-13B, Llama-2-70B, LLaMA-7B, LLaMA-13B, LLaMA-30B, Pythia-2.8B, Pythia-6.9B, Pythia-12B · paper: Function Vectors in Large Language Models, In-Context Learning Creates Task Vectors, Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning, Do Different Prompting Methods Yield a Common Task Representation in Language Models?
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B, Llama-3.1-8B, PaliGemma2 3B mix-448, PaliGemma2 10B mix-448, Idefics3-8B-Llama3 · paper: Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
structure: Curvature profile of the representation manifold, Linear Subspace · models: Llama-3.2-1B · paper: The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-2B-it · paper: Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
structure: Linear Direction · models: Gemma 3 4B Instruct, Llama-3.2-3B-Instruct · paper: Sycophancy Hides Linearly in the Attention Heads
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Qwen-2.5-7B, Qwen2.5-7B-Instruct, Gemma-2-9B, Gemma-2-9B-it, Mistral-7B-v0.3, Mistral-7B-Instruct-v0.3, OLMo 2 7B, OLMo 2 7B Instruct · paper: How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
structure: Linear Direction · models: Alpaca, Llama-3-8B-Lexi-Uncensored · paper: The Effectiveness of Style Vectors for Steering LLMs: A Human Evaluation
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-2-7B-Chat · paper: Probing then Editing Response Personality of Large Language Models
structure: Linear Direction · models: Geneformer · paper: Exhaustive Circuit Mapping of a Single-Cell Foundation Model Reveals Massive Redundancy, Heavy-Tailed Hub Architecture, and Layer-Dependent Differentiation Control
structure: Linear Representation Hypothesis · models: Llama-2-7B, Gemma-2B · paper: The Linear Representation Hypothesis and the Geometry of Large Language Models
structure: Linear Direction · models: Qwen3-8B, Qwen3-14B, Qwen3-32B, Gemma 3 4B Instruct, Gemma 3 12B Instruct, Llama-3.1-8B · paper: Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
structure: 1D continuum manifold, Linear Subspace · models: Llama-2-7B · paper: Probing for Representation Manifolds in Superposition
structure: Linear Direction, Linear Subspace · models: Gemma-2-2B-it, Gemma-2-9B-it, Llama-3.1-8B-Instruct · paper: SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
structure: Linear Direction · models: OthelloGPT · paper: Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT
structure: Linear Direction · models: OthelloGPT · paper: Emergent Linear Representations in World Models of Self-Supervised Sequence Models
structure: Linear Direction · models: DiffRhythm, EnCodec, WavTokenizer, Stable Audio Open · paper: Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
structure: Linear Direction · models: Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Steering at the Source: Style Modulation Heads for Robust Persona Control
structure: Linear Direction · models: DeepSeek-R1-Distill-Qwen-1.5B, DeepSeek-R1-Distill-Qwen-14B, DeepSeek-R1-Distill-Llama-8B · paper: Understanding Reasoning in Thinking Language Models via Steering Vectors
structure: Linear Direction · models: Llama-3-70B-Instruct, Gemma-7B-it, Qwen-1.8B-Chat, Vicuna-13B, Qwen1.5-1.8B-Chat, Qwen1.5-32B-Chat, Llama-2-13B-Chat, Llama-3.1-8B-Instruct, NeuralDaredevil-8B-abliterated, Hermes-2-Pro-Llama-3-8B, OLMo-7B-SFT, Zephyr-7B-Beta, H2O-Danube3-4B-Chat, Gemma-2-9B-it, Qwen2.5-7B-Instruct · paper: Refusal in Language Models Is Mediated by a Single Direction, Representation Engineering: A Top-Down Approach to AI Transparency, Programming Refusal with Conditional Activation Steering, Refusal Direction is Universal Across Safety-Aligned Languages
structure: Linear Direction · models: Llama-3-8B, Aya-23-8B · paper: Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
structure: Linear Subspace · models: Stable Diffusion v2.1 · paper: Be Tangential to Manifold: Discovering Riemannian Metric for Diffusion Models
structure: Linear Direction · models: Stable Diffusion v1.4 · paper: Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
structure: Linear Direction · models: GPT-2 Small, Pythia-1.4B, Pythia-2.8B · paper: Linear Representations of Sentiment in Large Language Models
structure: Linear Direction · models: Qwen2.5-14B-Instruct · paper: Convergent Linear Representations of Emergent Misalignment
structure: Linear Direction, Concept Crystals (Parallelogram/Trapezoid Structure) · models: Llama-3.2-3B-Instruct, Llama-3.2-1B-Instruct, Qwen3-1.7B · paper: Linear Spatial World Models Emerge in Large Language Models
structure: Linear Direction · models: ESM-2 (3B) · paper: Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
structure: Linear Subspace, Linear Direction · models: Gemma-2-2B-it, Llama-2-7B-Chat · paper: The Cylindrical Representation Hypothesis for Language Model Steering
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Qwen2.5-3B-Instruct · paper: On the Non-Identifiability of Steering Vectors in Large Language Models
structure: Linear Direction · models: Llama-3.2-1B-Instruct, Qwen2.5-0.5B-Instruct, Gemma-3-1B-it, Llama-3.1-8B-Instruct, Llama-3-8B-Instruct, Gemma-3-270M-it · paper: Steered LLM Activations Are Non-Surjective
structure: Linear Direction · models: Qwen3-30B-A3B-Instruct, Qwen3-4B-Instruct, Llama-3.1-8B-Instruct, Llama 3.3 70B Instruct, GPT-OSS-20B · paper: Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
structure: Linear Subspace · models: Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.1 · paper: Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention
structure: Linear Separability, Linear Direction · models: Mistral Small 3 (2501), Mistral-7B, Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-3.2-3B, Gemma-2-9B, Gemma-2-2B, GPT-J-6B, GPT-2 XL, GPT-2 Large, GPT-2 Medium, GPT-2 Small · paper: Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
structure: Circle, Cone, Platonic Representation Hypothesis · models: GPT-2-small, Mistral-7B, Llama-3-8B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Llama-3.1-8B · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, Not All Language Model Features Are One-Dimensionally Linear, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Do Sparse Autoencoders Capture Concept Manifolds?
structure: Linear Direction · models: LLaMA-7B, Alpaca-7B, Vicuna-7B · paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
structure: Circle, Linear Direction · models: Llama-3.1-8B-Instruct, Qwen3-8B, Qwen3-14B, Apertus-8B-Instruct-2509, Gemma 4 E4B-it · paper: Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control, Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
structure: Linear Direction · models: Chronos, MOMENT, Moirai-1.1-R Large · paper: Exploring Representations and Interventions in Time Series Foundation Models
structure: Linear Direction · models: DrugAssist, GeLLM3O-LLaMA3, GeLLM3O-Mistral, MolGen · paper: SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing
structure: Linear Separability, Linear Direction · models: Mistral-7B-Instruct, DeepSeek-LLM-7B-Chat · paper: Language Models Represent Beliefs of Self and Others