Extracts a candidate feature direction as the difference between mean activations of two contrastive prompt sets (e.g. harmful vs. harmless instructions) at a chosen layer and token position — no training required.
Used in (63 observations)
structure: Linear Direction · models: Qwen3-8B, GPT-OSS-20B, Qwen2.5-7B-Instruct · paper: What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
structure: Linear Subspace, Linear Direction · models: Gemma 2 27B Instruct, Gemma-2-27B, Qwen3-32B, Llama 3.3 70B Instruct · paper: The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
structure: Linear Subspace, Linear Direction · models: Qwen3-8B, Llama-3.1-8B-Instruct · paper: The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models
structure: Linear Direction · models: Qwen3-VL-4B-Instruct, Qwen3-VL-8B-Instruct, LLaVA-OneVision-1.5-4B-Instruct · paper: Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
structure: Linear Direction, Linear Subspace · models: Gemma 3 4B Instruct, Gemma 3 4B, Gemma 3 27B Instruct, Qwen3-4B-Instruct, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Tool Calling Is Linearly Readable and Steerable in Language Models
structure: Linear Direction · models: Llama-3.2-3B-Instruct, Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, Gemma 3 4B Instruct, Qwen2.5-7B-Instruct · paper: Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
structure: Linear Direction · models: Llama-3.1-8B, Qwen2.5-14B, Qwen-2.5-7B, Qwen2.5-0.5B, Llama-3.2-1B · paper: Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
structure: Linear Direction · models: Whisper small, Whisper large-v3 · paper: Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
structure: Linear Direction · models: BERT-base-uncased, GPT-2-small · paper: Investigating Aspect Features in Contextualized Embeddings with Semantic Scales and Distributional Similarity
structure: Linear Direction · models: Gemma 3 12B Instruct, Llama-3.1-8B-Instruct · paper: Dissociating the Internal Representations of Sycophancy in LLMs
structure: Linear Direction, Linear Separability · models: Qwen2.5-Coder-14B-Instruct, Llama-3.1-8B-Instruct, OLMo-7B · paper: Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization
structure: Conceptual Belief Space Hypothesis, Linear Subspace · models: Llama-3.1-8B-Instruct · paper: Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
structure: Linear Direction · models: Llama-3-8B-Instruct, Ministral-8B-Instruct-2410 · paper: Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
structure: Linear Direction · models: TabPFN v2 (tabular in-context-learning foundation model), Mitra, TabICLv2 · paper: A Mechanistic Study of Tabular Foundation Models
structure: Linear Direction · models: Llama-2-7B-Chat · paper: Understanding (Un)Reliability of Steering Vectors in Language Models
structure: Linear Direction · models: Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Persona Vectors: Monitoring and Controlling Character Traits in Language Models
structure: Linear Direction, Linear Separability · models: Llama-3.2-1B, Llama-3.2-3B, Gemma-2-2B, Qwen2.5-1.5B, Llama-3.1-8B, Mistral-7B-v0.3 · paper: Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations
structure: Linear Direction · models: ChessGPT-8L-25M, ChessGPT-16L-50M · paper: Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
structure: Linear Direction · models: wav2vec 2.0 Large (LV-60), HuBERT Large (LibriLight 60k), WavLM Large · paper: [b]=[d]-[t]+[p]: Self-Supervised Speech Models Discover Phonological Vector Arithmetic
structure: Linear Direction · models: Claude Sonnet 4.5 · paper: Emotion Concepts and their Function in a Large Language Model
structure: Linear Subspace, Linear Direction · models: GPT-2 Small, GPT-2 Medium, GPT-2 Large, Qwen2-1.5B-Instruct, Qwen2-7B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Gemma-2-2B-it · paper: The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
structure: Linear Direction · models: Qwen2.5-3B, Gemma 3 4B Instruct · paper: Unveiling the Latent Directions of Reflection in Large Language Models
structure: Linear Direction · models: Llama-3.1-8B, Mistral-7B-v0.3, Qwen2.5-7B-Instruct · paper: Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs
structure: Linear Subspace · models: CosyVoice2 · paper: A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
structure: Linear Direction, Linear Subspace · models: Llama-3.2-3B-Instruct, Qwen3-4B-Instruct, Gemma 3 4B Instruct · paper: Scenario-based Probing and Steering Cultural Values in Large Language Models
structure: Linear Direction · models: Qwen3.5-4B, Qwen3.5-4B-Instruct, Llama-3.2-3B-Instruct, Qwen3.6-27B-Instruct, Gemma-4-31B-Instruct · paper: GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
structure: Linear Direction · models: Qwen2.5-VL 7B Instruct, Qwen3-VL 8B, Qwen3-VL 32B, LLaVA-OneVision-1.5 8B, Gemma 3 4B Instruct, Gemma 3 27B Instruct · paper: Causal Probing for Internal Visual Representations in Multimodal Large Language Models
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B, Llama-3.1-8B, PaliGemma2 3B mix-448, PaliGemma2 10B mix-448, Idefics3-8B-Llama3 · paper: Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Qwen-2.5-7B, Qwen2.5-7B-Instruct, Gemma-2-9B, Gemma-2-9B-it, Mistral-7B-v0.3, Mistral-7B-Instruct-v0.3, OLMo 2 7B, OLMo 2 7B Instruct · paper: How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
structure: Linear Direction · models: DDPM (CelebA-HQ 256x256, HuggingFace Diffusers), DDPM (LSUN-Church 256x256, HuggingFace Diffusers), DDPM (LSUN-Bedroom 256x256, HuggingFace Diffusers) · paper: Discovering Interpretable Directions in the Semantic Latent Space of Diffusion Models
structure: Linear Direction · models: Qwen2.5-7B, Llama-3.1-8B-Instruct, Gemma-2-9B-it · paper: Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
structure: Linear Direction · models: LLaVA-1.5-7B, Llama-3.1-8B-Instruct, InternVL3-8B, Gemma-7B-it · paper: Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
structure: Linear Direction · models: AudioLDM2, Stable Audio Open, Ace-Step · paper: TADA! Tuning Audio Diffusion Models Through Activation Steering
structure: Linear Representation Hypothesis · models: Llama-2-7B, Gemma-2B · paper: The Linear Representation Hypothesis and the Geometry of Large Language Models
structure: Linear Subspace · models: Llama-3.2-3B, Llama-3-8B · paper: Do Linear Probes Generalize Better in Persona Coordinates?
structure: Linear Direction · models: Molmo-7B-O, NVILA-Lite-2B, Qwen2.5-VL-3B-Instruct, RoboRefer-2B-SFT, Qwen3-VL-235B-A22B-Instruct · paper: Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
structure: Linear Direction · models: Qwen3-1.7B, Llama-3.1-8B-Instruct, Gemma 3 1B IT, Qwen2.5-7B-Instruct · paper: Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
structure: Linear Direction · models: Gemma 2 9B IT, Gemma 3 12B IT, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct · paper: A Retrieval-Conditioned Rebinding Circuit for Dynamic Entity Tracking in Large Language Models
structure: Linear Separability, Linear Direction · models: HuBERT-base · paper: Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations
structure: Linear Direction · models: DeepSeek-R1-Distill-Qwen-1.5B, DeepSeek-R1-Distill-Qwen-14B, DeepSeek-R1-Distill-Llama-8B · paper: Understanding Reasoning in Thinking Language Models via Steering Vectors
structure: Affine Subspace · models: Llama-3-8B-Instruct, Llama-3-70B-Instruct, Hermes Eagle RWKV v5 (7B) · paper: Refusal in LLMs is an Affine Function
structure: Cone · models: Gemma-2-2B-it, Gemma-2-9B-it, Qwen2.5-1.5B-Instruct, Qwen2.5-14B-Instruct, Llama-3-8B-Instruct · paper: The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
structure: Linear Direction · models: Gemma-2-9B-it, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct · paper: Refusal Direction is Universal Across Safety-Aligned Languages
structure: Linear Direction · models: Llama-3-70B-Instruct, Gemma-7B-it, Qwen-1.8B-Chat, Vicuna-13B, Qwen1.5-1.8B-Chat, Qwen1.5-32B-Chat, Llama-2-13B-Chat, Llama-3.1-8B-Instruct, NeuralDaredevil-8B-abliterated, Hermes-2-Pro-Llama-3-8B, OLMo-7B-SFT, Zephyr-7B-Beta, H2O-Danube3-4B-Chat, Gemma-2-9B-it, Qwen2.5-7B-Instruct · paper: Refusal in Language Models Is Mediated by a Single Direction, Representation Engineering: A Top-Down Approach to AI Transparency, Programming Refusal with Conditional Activation Steering, Refusal Direction is Universal Across Safety-Aligned Languages
structure: Cone · models: Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen2.5-7B-Instruct · paper: Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
structure: Linear Direction · models: GPT-2 Small, Pythia-1.4B, Pythia-2.8B · paper: Linear Representations of Sentiment in Large Language Models
structure: Linear Direction · models: Qwen2.5-14B-Instruct · paper: Convergent Linear Representations of Emergent Misalignment
structure: Linear Direction, Linear Subspace · models: wav2vec 2.0 Large (LV-60), HuBERT Large (LibriLight 60k), WavLM Large · paper: Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
structure: Linear Subspace, Linear Direction · models: Gemma-2-2B-it, Llama-2-7B-Chat · paper: The Cylindrical Representation Hypothesis for Language Model Steering
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-2-2B-it · paper: ReCoVeR the Target Language: Language Steering Without Sacrificing Task Performance
structure: Linear Direction · models: Qwen3-30B-A3B-Instruct, Qwen3-4B-Instruct, Llama-3.1-8B-Instruct, Llama 3.3 70B Instruct, GPT-OSS-20B · paper: Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
structure: Linear Direction · models: Qwen2.5-1.5B-Instruct, Gemma-2-2B-it, Llama-3.2-1B-Instruct, Ministral-3-3B-Instruct · paper: Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families
structure: Linear Direction · models: Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-2-9B-it, Gemma-2-2B-it · paper: The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
structure: Linear Separability · models: Qwen 0.5B (base), GPT-2 XL, Qwen 7B (base), Llama-2-7B, Llama-2-7B-Chat, Gemma 1.1 7B Instruct, Qwen2-7B-Instruct · paper: The Blessing and Curse of Dimensionality in Safety Alignment
structure: Cone · models: Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Gemma-2-2B-it, Gemma-2-9B-it · paper: From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
structure: Linear Direction · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-v0.1 · paper: The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets, On the Universal Truthfulness Hyperplane Inside LLMs
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Mistral NeMo 12B Instruct (2407), Qwen3-4B-Instruct, SmolLM3-3B · paper: How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
structure: Linear Direction · models: LLaMA-7B, Alpaca-7B, Vicuna-7B · paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
structure: Linear Direction · models: Llama-3.2-1B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct · paper: Negative Before Positive: Asymmetric Valence Processing in Large Language Models
structure: Linear Direction · models: Chronos, MOMENT, Moirai-1.1-R Large · paper: Exploring Representations and Interventions in Time Series Foundation Models
structure: Linear Direction · models: Qwen3-32B, Llama 3.3 70B Instruct · paper: Rhetorical Questions in LLM Representations: A Linear Probing Study
structure: Linear Direction · models: Llama-3-8B-Instruct, Llama-2-7B-Chat · paper: Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
structure: Linear Direction · models: · paper: Characterising Universal Jailbreak Features and Refusal Direction in LLMs