Directly adds or ablates a fixed direction in activation space via vector arithmetic and observes the effect on model output — no second forward pass needed, unlike activation patching.
Used in (153 observations)
structure: Linear Subspace · models: Llama-3.1-8B-Instruct, Qwen2-7B-Instruct · paper: Death by a Thousand Directions: Exploring the Geometry of Harmfulness in LLMs through Subconcept Probing
structure: Lissajous Curves · models: Modular-Arithmetic Vanilla RNN (1 layer, tanh, d_h=256, mod 113) · paper: Modular Addition in Recurrent Neural Networks Requires Low-Rank Fourier Circuits
structure: Linear Subspace · models: Gemma-2-2B-it, Llama-3.1-8B-Instruct · paper: Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
structure: Linear Direction, Linear Separability · models: Llama 3.3 70B Instruct, Qwen2.5-72B · paper: Understanding Moral Reasoning Trajectories in LLMs: Toward Probing-Based Explainability
structure: Linear Direction, Linear Separability · models: Llama 3.3 70B Instruct · paper: Probing and Steering Evaluation Awareness of Language Models
structure: Linear Direction · models: Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Gemma-2-2B, Llama-3.2-3B, SmolLM3-3B, Qwen3-0.6B · paper: A Circuit for Predicting Hierarchical Structure In-Context in Large Language Models
structure: Linear Subspace · models: Llama-3-8B-Instruct, Qwen3-8B · paper: Cell-Based Representation of Relational Binding in Language Models
structure: Linear Subspace, Linear Direction · models: Llama-2-7B, Llama-3-8B, Qwen1.5-7B, Pythia-6.9B, Float-7B (code fine-tuned LM) · paper: Representational Analysis of Binding in Language Models
structure: Linear Subspace · models: ResNet-18 (supervised, ImageNet), ResNet-18 (SimCLR contrastive pretraining, CIFAR-10) · paper: Simple Disentanglement of Style and Content in Visual Representations
structure: Line Attractor · models: Piecewise-linear RNN (N=40, delayed-addition short-term-memory task) · paper: Understanding and Controlling the Geometry of Memory Organization in RNNs
structure: Linear Direction · models: IndexTTS2 · paper: Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
structure: Linear Direction · models: AlphaZero (chess, ResNet policy/value network, self-play) · paper: Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
structure: Linear Direction · models: Llama-3.2-3B-Instruct, Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, Gemma 3 4B Instruct, Qwen2.5-7B-Instruct · paper: Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
structure: Linear Direction · models: Llama-3.1-8B, Qwen2.5-14B, Qwen-2.5-7B, Qwen2.5-0.5B, Llama-3.2-1B · paper: Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
structure: Torus · models: LSTM path-integration module + A3C navigation agent · paper: Vector-based Navigation Using Grid-like Representations in Artificial Agents
structure: Linear Subspace, Linear Direction · models: Bayesian Wind Tunnel Transformer (bijection task, 6 layers, 6 heads, d_model=192), Bayesian Wind Tunnel Transformer (HMM filtering task, 9 layers, 8 heads, d_model=256) · paper: The Bayesian Geometry of Transformer Attention
structure: Belief State Geometry Hypothesis (Mixed-State Presentation), Polytope (Simplex) · models: Gemma-2-9B · paper: Finding Belief Geometries with Sparse Autoencoders
structure: Linear Direction · models: β-VAE (convolutional encoder/decoder, dSprites/colored-dSprites/CelebA/3D-Chairs) · paper: Understanding disentangling in β-VAE
structure: Linear Direction · models: TabPFN v2 (tabular in-context-learning foundation model), Mitra, TabICLv2 · paper: A Mechanistic Study of Tabular Foundation Models
structure: Linear Direction · models: word2vec (Google News, 300d) · paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
structure: Linear Separability · models: InternVL3-8B · paper: Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection
structure: Linear Direction · models: Gemma-2-2B, Qwen2-0.5B, Llama-3.2-1B · paper: A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
structure: Linear Direction · models: 128-dim dot-product matrix-factorization recommender, trained on Alibaba/Tianchi mobile-recommendation logs (Cheng) · paper: A Rank-One Popularity Component in Dot-Product Recommender Scores: Population Theory and Prior-Separation Evidence
structure: Circle · models: Small MLP/Transformer trained on finite-group composition tasks · paper: A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
structure: Linear Direction · models: scGPT (whole-human pretrained checkpoint), scFoundation (pretrained) · paper: Sparse Autoencoders Reveal Interpretable Features in Single-Cell Foundation Models
structure: Linear Direction · models: Claude Sonnet 4.5 · paper: Emotion Concepts and their Function in a Large Language Model
structure: Anisotropy, Linear Direction · models: CLIP ViT-B/32, CLIP ViT-B/16 · paper: Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
structure: Dimensional collapse · models: CLIP ViT-B/16, CLIP ViT-B/32, CLIP ViT-L/14, SigLIP, SigLIP 2 · paper: Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
structure: Linear Subspace, Anisotropy · models: CLAP (HTSAT-BERT-ZS, trained on WavCaps), DRCap's CLAP (trained on WavCaps + SoundVECaps) · paper: COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings
structure: Linear Direction · models: Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B, Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Qwen-2.5-7B, Qwen2.5-14B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Gemma-2-2B, Gemma-2-9B, Mistral-7B-v0.3, SmolLM2-1.7B, OLMoE-1B-7B, JetMoE-8B · paper: Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations
structure: Linear Direction · models: Llama-2-7B-Chat, Llama-2-13B-Chat, Llama-2-70B-Chat, Vicuna-13B, Vicuna-33B-Uncensored, DeBERTa-xxlarge-v2-MNLI · paper: Representation Engineering: A Top-Down Approach to AI Transparency
structure: Torus · models: Conformal-normalization linear path-integration RNN, Conformal-normalization nonlinear (CANN-like) path-integration RNN · paper: Emergence of Grid-like Representations by Training Recurrent Networks with Conformal Normalization
structure: Linear Direction · models: Qwen2.5-3B, Gemma 3 4B Instruct · paper: Unveiling the Latent Directions of Reflection in Large Language Models
structure: Linear Subspace · models: CosyVoice2 · paper: A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
structure: Linear Subspace · models: XLM-RoBERTa base · paper: The Geometry of Multilingual Language Model Representations
structure: Linear Direction, Anisotropy · models: CLIP ViT-B/32, MetaCLIP ViT-B/32, OpenCLIP ViT-B/32, SigLIP 2 · paper: Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
structure: Lissajous Curves · models: OLMo 2 1B, OLMo 2 7B, OLMo 2 13B, Llama-3.2-1B, Llama-3.2-3B, Llama-3-8B, Llama-3.1-8B, Phi-4 (15B) · paper: Unravelling the Mechanisms of Manipulating Numbers in Language Models
structure: Linear Direction, Linear Subspace · models: Llama-3.2-3B-Instruct, Qwen3-4B-Instruct, Gemma 3 4B Instruct · paper: Scenario-based Probing and Steering Cultural Values in Large Language Models
structure: Linear Direction · models: Qwen2.5-VL 7B Instruct, Qwen3-VL 8B, Qwen3-VL 32B, LLaVA-OneVision-1.5 8B, Gemma 3 4B Instruct, Gemma 3 27B Instruct · paper: Causal Probing for Internal Visual Representations in Multimodal Large Language Models
structure: Linear Subspace, Linear Direction · models: Qwen3-4B, Qwen3-8B, OLMo 3 7B · paper: Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics
structure: Linear Direction · models: DRC(3,3) Sokoban Agent · paper: Interpreting Emergent Planning in Model-Free Reinforcement Learning
structure: Anisotropy · models: LLaVA-1.5-7B, LLaVA-1.5-13B, InternVL3-8B · paper: What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
structure: Circle · models: Grokking Modular-Arithmetic Transformer (2 layers, 4 heads, pre-LN, d_model=128, mod 113/149/197) · paper: Circuit Synchronization Precedes Generalization: A Causal Precursor to Grokking
structure: Linear Direction · models: OpenFlamingo-4B · paper: Multimodal Function Vectors for Visual Relations
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B, Gemma-2-27B, Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B-Instruct · paper: How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
structure: Linear Subspace, Linear Direction · models: Leela Chess Zero (policy network) · paper: The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network
structure: Linear Direction · models: StyleGAN2 (trained on FFHQ, 1024x1024), BigGAN-deep (512px) · paper: GANSpace: Discovering Interpretable GAN Controls
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-2B-it · paper: A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
structure: Lissajous Curves · models: GPT-2-XL · paper: Pre-trained Large Language Models Use Fourier Features to Compute Addition
structure: Intrinsic-dimension profile across depth · models: Stable Diffusion XL, SDXL-DMD (4-step distilled) · paper: ELROND: Exploring and Decomposing Intrinsic Capabilities of Diffusion Models
structure: Linear Direction · models: Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B · paper: Latent Planning Emerges with Scale
structure: Linear Direction · models: OpenVLA-7B, pi0 (PaliGemma VLA backbone, predecessor of pi0.5) · paper: Mechanistic Interpretability for Steering Vision-Language-Action Models
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Llama 3.3 70B Instruct, Qwen3-0.6B, Qwen3-4B, Gemma 3 4B Instruct, Gemma 3 12B Instruct · paper: Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
structure: Decision boundary (as a codimension-1 hypersurface), Linear Direction · models: Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, Gemma-2-9B-it, Qwen2.5-7B-Instruct, Phi-3.5-mini-instruct, Llama-3-8B · paper: Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
structure: Circle, 2D Grid (Square Lattice) · models: Llama-3.1-8B, Llama-3.2-1B, Llama-3.1-8B-Instruct, Gemma-2-2B, Gemma-2-9B · paper: In-Context Learning of Representations
structure: Linear Direction · models: Llama-3-8B-Instruct, Pythia-70M · paper: Inference-Time Causal Probing in LLMs
structure: Linear Direction · models: PGGAN (trained on CelebA-HQ, 1024x1024), StyleGAN (trained on FFHQ, 1024x1024) · paper: InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs
structure: Linear Subspace, Anisotropy · models: CLIP ViT-B/32, CLIP ViT-L/14, OpenCLIP ViT-B/32, OpenCLIP ViT-L/14, SigLIP, SigLIP 2 · paper: Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
structure: Linear Direction · models: OpenCLIP ViT-B/16 (LAION-2B), DINOv2 ViT-L/14, LLaVA-Llama-3-8B · paper: Vision Transformers Don't Need Trained Registers
structure: Anisotropy · models: Llama-3-8B-Instruct, Gemma-2-9B-it, Qwen2.5-7B-Instruct, Llama-2-7B-Chat, GPT-2-small, OPT-350M, OPT-2.7B · paper: Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qwen2.5-7B · paper: Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
structure: Intrinsic-dimension profile across depth · models: GPT-2 XL, Pythia-2.8B, Custom 12-layer GPT-2-Small (trained from scratch, 100M tokens) · paper: Representational Curvature Modulates Behavioral Uncertainty in Large Language Models
structure: Linear Direction · models: CosyVoice3 (Qwen2.5-0.5B backbone) · paper: Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
structure: Linear Direction · models: XLM-RoBERTa base · paper: The Geometry of Multilingual Language Model Representations
structure: Intrinsic-dimension profile across depth · models: OPT-125M, OPT-1.3B, OPT-13B, Pythia-160M, Pythia-410M, Pythia-6.9B, WavLM-base-plus, WavLM Large, Whisper large · paper: Abstraction Induces the Brain Alignment of Language and Speech Models
structure: Nonlinear World-Model Decodability · models: OthelloGPT · paper: Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
structure: Linear Direction, Linear Subspace · models: BART-base, T5-Base · paper: Implicit Representations of Meaning in Neural Language Models
structure: Linear Separability · models: Ovis2.5-2B, InternVL3.5-2B, VST-3B · paper: Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
structure: Circle · models: pGESAM (pitch-conditioned Generative Sample Map), EnCodec · paper: Pitch-Conditioned Instrument Sound Synthesis from an Interactive Timbre Latent Space
structure: Linear Separability, Linear Direction · models: Llama-2-7B-Chat, Vicuna-7B · paper: Unlocking the Future: Look-Ahead Planning Mechanistic Interpretability in LLMs
structure: Circle · models: Llama-3-8B, Mistral-7B · paper: Language Models Encode Numbers Using Digit Representations in Base 10
structure: Linear Subspace · models: Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct · paper: From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task
structure: Anisotropy · models: LLaVA-1.5-7B · paper: Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
structure: Linear Subspace, Linear Direction · models: CodeLlama-13B, Gemma-2-2B, Llama 3.1 70B · paper: Do Language Models Track Entities Across State Changes?
structure: Linear Direction · models: AudioLDM2, Stable Audio Open, Ace-Step · paper: TADA! Tuning Audio Diffusion Models Through Activation Steering
structure: 1D continuum manifold · models: Pythia-2.8B, Llama-2-7B, Llama-3.1-8B, Llama-3.2-1B, GPT-2-Large, Mistral-7B, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Gemma-2-2B, Llama-3.1-70B-Instruct, Llama-3-8B-Instruct, Llama-3-8B, Mistral-7B-Instruct-v0.3, Qwen2.5-7B-Instruct · paper: Number Representations in LLMs: A Computational Parallel to Human Perception, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models, LLMs Know More About Numbers than They Can Say
structure: Linear Direction · models: Gemma-2-2B, Llama-3.2-3B-Instruct, GPT-2-XL · paper: Universal Response and Emergence of Induction in LLMs
structure: Linear Direction · models: Llama 3.3 70B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-7B-Instruct · paper: The Truthfulness Spectrum Hypothesis
structure: Linear Subspace · models: GPT-2 Small, GPT-2 Medium, Qwen2.5-Math-1.5B · paper: Invariant Reasoning Directions in Latent Trajectories of Language Models
structure: Anisotropy · models: ViT-S/16 (12 layers, 8 heads, dim 384, trained from scratch on ImageNet-100) · paper: Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
structure: Dimensional collapse · models: Pythia-410M, Pythia-6.9B, GPT-OSS-20B, Gemma-7B · paper: Attention Sinks and Compression Valleys in LLMs Are Two Sides of the Same Coin
structure: Linear Direction · models: FLUX.1 [schnell] · paper: DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing
structure: Linear Direction · models: GPT-J-6B, GPT-2 Small, GPT-2 Medium, GPT-2 Large, GPT-2 XL, BLOOM-176B · paper: Language Models Implement Simple Word2Vec-style Vector Arithmetic
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-2B-it · paper: Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
structure: Linear Direction · models: Qwen3-1.7B, Llama-3.1-8B-Instruct, Gemma 3 1B IT, Qwen2.5-7B-Instruct · paper: Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
structure: Torus, Circle · models: Grokking Modular-Arithmetic Transformer (1 layer, 4 heads, d_model=128, mod 113), Clock/Pizza Transformer, Model A (1 layer, constant attention alpha=0, width 128, mod 59), Clock/Pizza Transformer, Model B (1 layer, normal attention alpha=1, width 128, mod 59) · paper: Progress Measures for Grokking via Mechanistic Interpretability, The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks, On the Geometry and Topology of Representations: The Manifolds of Modular Addition
structure: Anisotropy · models: word2vec (Google News, 300d), GloVe (840B token Common Crawl) · paper: All-but-the-Top: Simple and Effective Postprocessing for Word Representations
structure: Linear Direction · models: Llama-3.1-8B, Mistral-7B-v0.1 · paper: How Language Models Process Negation
structure: Linear Direction · models: Phi-3-mini, Llama-3.1-8B-Instruct · paper: Steering Language Model Refusal with Sparse Autoencoder Features
structure: Linear Direction · models: Gemma 2 27B Instruct, Qwen3-32B · paper: Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
structure: Linear Direction · models: Gemma 2 9B IT, Gemma 3 12B IT, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct · paper: A Retrieval-Conditioned Rebinding Circuit for Dynamic Entity Tracking in Large Language Models
structure: Linear Direction · models: text-embedding-3-small · paper: Disentangling Dense Embeddings with Sparse Autoencoders
structure: Linear Direction · models: DiffRhythm, EnCodec, WavTokenizer, Stable Audio Open · paper: Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
structure: Intrinsic-dimension profile across depth · models: TS-51M (custom 8-layer x 512d x 16-head transformer, TinyStories, 6 pretraining seeds), nanoGPT-style GPT-2 124M (FineWeb-10B), Pythia-160M, Pythia-410M, Pythia-1B, OLMo-1B, OLMoE-1B-7B · paper: Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
structure: Linear Subspace · models: Custom GPT-2-Small (trained from scratch on synthetic grid-navigation token sequences) · paper: Cognitive Maps in Language Models: A Mechanistic Analysis of Spatial Planning
structure: Linear Direction · models: Whisper base · paper: Mechanistic Interpretability of ASR models using Sparse Autoencoders
structure: Linear Subspace · models: CLIP ViT-B/32 · paper: Parts of Speech-Grounded Subspaces in Vision-Language Models
structure: Linear Subspace · models: ELMo (5.5B-word pretrained, 2-layer biLSTM), BERT-base-cased · paper: The Low-Dimensional Linear Geometry of Contextualized Word Representations
structure: Linear Direction · models: ProSoRo Multi-modal Variational Autoencoder (motion/force/shape latent proprioception) · paper: Anchoring Morphological Representations Unlocks Latent Proprioception in Soft Robots
structure: Linear Direction · models: DCGAN (trained on aligned & cropped celebrity faces) · paper: Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
structure: Linear Direction · models: Byte mLSTM (4096-unit, trained on ~82M Amazon reviews) · paper: Learning to Generate Reviews and Discovering Sentiment
structure: Affine Subspace · models: Llama-3-8B-Instruct, Llama-3-70B-Instruct, Hermes Eagle RWKV v5 (7B) · paper: Refusal in LLMs is an Affine Function
structure: Cone · models: Gemma-2-2B-it, Gemma-2-9B-it, Qwen2.5-1.5B-Instruct, Qwen2.5-14B-Instruct, Llama-3-8B-Instruct · paper: The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
structure: Linear Direction · models: Gemma-2-9B-it, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct · paper: Refusal Direction is Universal Across Safety-Aligned Languages
structure: Linear Direction · models: Llama-3-70B-Instruct, Gemma-7B-it, Qwen-1.8B-Chat, Vicuna-13B, Qwen1.5-1.8B-Chat, Qwen1.5-32B-Chat, Llama-2-13B-Chat, Llama-3.1-8B-Instruct, NeuralDaredevil-8B-abliterated, Hermes-2-Pro-Llama-3-8B, OLMo-7B-SFT, Zephyr-7B-Beta, H2O-Danube3-4B-Chat, Gemma-2-9B-it, Qwen2.5-7B-Instruct · paper: Refusal in Language Models Is Mediated by a Single Direction, Representation Engineering: A Top-Down Approach to AI Transparency, Programming Refusal with Conditional Activation Steering, Refusal Direction is Universal Across Safety-Aligned Languages
structure: Linear Direction · models: Llama-2-7B, GPT-J-6B · paper: Identifying Linear Relational Concepts in Large Language Models
structure: Linear Direction · models: MusicGen-Large · paper: Steering Autoregressive Music Generation with Recursive Feature Machines
structure: Cone · models: Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen2.5-7B-Instruct · paper: Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
structure: Anisotropy · models: BERT-base-cased, RoBERTa-base, GPT-2-small, XLNet-base-cased, word2vec (Google News, 300d) · paper: All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality
structure: Linear Direction · models: Llama-3.1-8B, Llama 3.1 70B · paper: Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction
structure: Linear Direction · models: Llama-3-8B, Aya-23-8B · paper: Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
structure: Linear Direction · models: Stable Diffusion v1.4 · paper: Emergence and Evolution of Interpretable Concepts in Diffusion Models
structure: Linear Subspace · models: Llama-3.2-3B, Llama-3.1-8B, Qwen3-8B, Qwen3-14B · paper: Linear Representations of Hierarchical Concepts in Language Models
structure: Linear Direction · models: ViT-B/16 (ImageNet-21k, supervised) · paper: From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers
structure: Linear Subspace · models: GPT-2-small, GPT-2-Medium, GPT-2-Large · paper: Multilinguality as Sense Adaptation
structure: Linear Direction, Linear Separability · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Llama-3-8B, Llama-3-70B, Gemma-2B, Gemma-7B · paper: Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
structure: Linear Subspace · models: E5-large-v2 · paper: Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
structure: Continuous causal attention operator (CCT) · models: Llama-2-13B-Chat, Llama-3-8B, Phi-3-Medium-4k-Instruct, Gemma-7B, Gemma-2-9B, Mistral-7B, GPT-2 · paper: Language Models Are Implicitly Continuous
structure: Linear Direction · models: DRC(3,3) Sokoban Agent · paper: Planning in a Recurrent Neural Network That Plays Sokoban
structure: Linear Direction · models: Qwen2.5-14B-Instruct · paper: Convergent Linear Representations of Emergent Misalignment
structure: Linear Direction, Anisotropy · models: OpenCLIP ViT-B/32, CLIP ResNet-50 · paper: Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-2-2B-it · paper: ReCoVeR the Target Language: Language Steering Without Sacrificing Task Performance
structure: Affine Subspace, Linear Subspace · models: StyleGAN2 (trained on FFHQ, 1024x1024) · paper: Do Not Escape From the Manifold: Discovering the Local Coordinates on the Latent Space of GANs
structure: Linear Separability · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, DeepSeek-R1-Distill-Llama-8B · paper: LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
structure: Linear Subspace · models: DINOv2 ViT-L/14, MAE ViT-Large (Masked Autoencoder), iBOT ViT-L/16 · paper: Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
structure: Dimensional collapse · models: Custom controlled 7B Transformer (trained from scratch for architecture/normalization ablations), Llama-2-7B · paper: The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
structure: Linear Direction · models: CLIP ViT-bigG/14 (Kandinsky 2.2 image encoder) · paper: Decoding Vision Transformers: The Diffusion Steering Lens
structure: Linear Direction · models: Mini-ICL synthetic RoPE transformer (6L/2H/d128 for dice and Markov-chain tasks; 16L for linear regression), Qwen-2.5-7B · paper: Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
structure: Linear Direction · models: Synthetic GPT-2-style ICL transformer (3-8 layers, trained from scratch on regression/token-offset/GINC/RegBench tasks) · paper: Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
structure: Linear Direction · models: Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-2-9B-it, Gemma-2-2B-it · paper: The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
structure: Linear Direction · models: Qwen3-4B-Instruct-2507 · paper: Temporal Preference Concepts and Their Functions in a Large Language Model
structure: Anisotropy, Linear Subspace · models: CLIP ViT-B/32 · paper: Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
structure: Anisotropy, Linear Direction · models: CLIP ViT-B/16 · paper: On the Modality Gap and the Contrastive Loss in Multi-modal Representation Learning
structure: Torus · models: CARNN trained on path integration across multiple (1-50) environments · paper: Coherently Remapping Toroidal Cells But Not Grid Cells are Responsible for Path Integration in Virtual Agents
structure: Limit cycle (stable periodic attractor) · models: GRU policy (PPO, Procgen Jumper), Mamba policy (PPO, Procgen Jumper) · paper: Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
structure: Jacobian non-normality (Schur decomposition of per-layer operators) · models: Llama-3.1-8B, Gemma 4 E4B, OLMo 3 7B · paper: Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
structure: Linear Direction · models: DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B · paper: Internal states before "wait" modulate reasoning patterns
structure: Cone · models: Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Gemma-2-2B-it, Gemma-2-9B-it · paper: From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
structure: Linear Direction · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-v0.1 · paper: The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets, On the Universal Truthfulness Hyperplane Inside LLMs
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Mistral NeMo 12B Instruct (2407), Qwen3-4B-Instruct, SmolLM3-3B · paper: How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
structure: Linear Direction · models: LLaMA-7B, Alpaca-7B, Vicuna-7B · paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
structure: Linear Direction · models: Llama-3.2-1B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct · paper: Negative Before Positive: Asymmetric Valence Processing in Large Language Models
structure: Linear Separability · models: Llama-3.2-1B, Gemma-2-2B · paper: Multilingual Language Models Encode Script Over Linguistic Structure
structure: Anisotropy · models: Mamba-130m (state-spaces/mamba-130m-hf), RoBERTa-base · paper: Lost in State Space: Probing Frozen Mamba Representations
structure: Stratified Manifold with Continuous Fibers · models: Qwen3-4B, Qwen3-8B, Gemma 3 4B Instruct · paper: The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models
structure: Linear Direction, Constraint-Algebra Basis Hypothesis · models: Sudoku Transformer (8 layers, 8 heads, d_model=576) · paper: Transformers Linearly Represent Highly Structured World Models
structure: Linear Direction · models: StyleNeRF (FFHQ, 256x256) · paper: NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation
structure: Linear Direction · models: FLUX.1, Stable Diffusion 3.5 · paper: Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
structure: Linear Direction · models: Qwen3.5-35B-A3B (MoE, ~3B active/token) · paper: Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits
structure: Linear Direction · models: Stable Diffusion v1.5 · paper: Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
structure: Linear Direction · models: Llama-3-8B-Instruct, Llama-2-7B-Chat · paper: Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
structure: Linear Direction · models: Pythia-70M, Pythia-160M, Pythia-410M, Pythia-1B, Pythia-1.4B, Pythia-2.8B, Pythia-6.9B, GPT-2-small, GPT-2-Medium, GPT-2-Large, GPT-2-XL, Llama-2-7B · paper: Which Attention Heads Matter for In-Context Learning?
structure: Linear Direction · models: IRIS (Atari 100k, Breakout/Pong), DIAMOND (Atari 100k, Breakout/Pong) · paper: What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators