Trains a sparse dictionary-learning autoencoder on model activations to recover a large set of candidate feature directions with no supervision; dictionary elements can be clustered to surface multi-dimensional features.
Used in (48 observations)
structure: 1D continuum manifold · models: Gemma-2-2B, Gemma-2-9B, Qwen3-4B · paper: Geometry of Ordinal Representations in Language Models, When Models Manipulate Manifolds: The Geometry of a Counting Task
structure: Linear Direction · models: Llama-3.1-8B · paper: Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
structure: Linear Subspace · models: Gemma-2-2B-it, Llama-3.1-8B-Instruct · paper: Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
structure: Linear Direction, Linear Subspace · models: Gemma 3 4B Instruct, Gemma 3 4B, Gemma 3 27B Instruct, Qwen3-4B-Instruct, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct · paper: Tool Calling Is Linearly Readable and Steerable in Language Models
structure: Linear Direction · models: IndexTTS2 · paper: Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
structure: Intrinsic-dimension profile across depth · models: Pythia-70M · paper: Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
structure: Dimensional collapse · models: Llama-3.1-8B, GPT-2, Gemma-2-9B, Qwen3-8B, Qwen3-4B, Pythia-160M, Pythia-2.8B · paper: Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
structure: Belief State Geometry Hypothesis (Mixed-State Presentation), Polytope (Simplex) · models: Gemma-2-9B · paper: Finding Belief Geometries with Sparse Autoencoders
structure: Linear Direction · models: Stable Diffusion v1.5 · paper: SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
structure: Linear Direction · models: Gemma-2-2B, Qwen2-0.5B, Llama-3.2-1B · paper: A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B · paper: Causal Language Control in Multilingual Transformers via Sparse Feature Steering
structure: Linear Direction · models: Claude 3 Sonnet · paper: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
structure: Linear Direction · models: FLUX.1 · paper: Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
structure: Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis) · models: DINOv2-B (ViT-Base, 4 register tokens) · paper: Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B, Llama-3.1-8B, PaliGemma2 3B mix-448, PaliGemma2 10B mix-448, Idefics3-8B-Llama3 · paper: Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-2B-it · paper: Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
structure: Intrinsic-dimension profile across depth · models: Gemma-2-2B, Gemma-2-9B · paper: The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
structure: Linear Direction · models: Pythia-70M, GPT-2-small, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Gemma-2-9B-it, Gemma-2-2B · paper: Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
structure: Linear Subspace · models: GPT-2 Small, Mistral-7B-v0.1 · paper: Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability
structure: Intrinsic-dimension profile across depth · models: Stable Diffusion XL, SDXL-DMD (4-step distilled) · paper: ELROND: Exploring and Decomposing Intrinsic Capabilities of Diffusion Models
structure: Linear Direction · models: HyenaDNA-small-32k · paper: Sparse Autoencoders Reveal Interpretable Structure in Small Gene Language Models
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Llama 3.3 70B Instruct, Qwen3-0.6B, Qwen3-4B, Gemma 3 4B Instruct, Gemma 3 12B Instruct · paper: Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
structure: Intrinsic-dimension profile across depth · models: ESM-2 (35M), ESM-2 (650M), ESM-2 (3B), iGPT-S, iGPT-M, iGPT-L, Llama-2-70B, GPT-2-XL, Gemma-2-2B · paper: The Geometry of Hidden Representations of Large Transformer Models, The Geometry of Concepts: Sparse Autoencoder Feature Structure
structure: Linear Direction · models: Geneformer · paper: Exhaustive Circuit Mapping of a Single-Cell Foundation Model Reveals Massive Redundancy, Heavy-Tailed Hub Architecture, and Layer-Dependent Differentiation Control
structure: Linear Direction · models: AudioLDM2, Stable Audio Open, Ace-Step · paper: TADA! Tuning Audio Diffusion Models Through Activation Steering
structure: Linear Direction · models: Qwen3-8B, Qwen3-14B, Qwen3-32B, Gemma 3 4B Instruct, Gemma 3 12B Instruct, Llama-3.1-8B · paper: Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
structure: Circle, 1D continuum manifold · models: GPT-2-small, Mistral-7B, text-embedding-3-large · paper: The Origins of Representation Manifolds in Large Language Models
structure: Linear Direction · models: Gemma-2-2B-it, Gemma-2-9B-it, Gemma 2 27B Instruct, Llama-3.1-8B-Instruct, GPT-OSS-20B · paper: Understanding Emergent Misalignment via Feature Superposition Geometry
structure: Linear Direction, Linear Subspace · models: Gemma-2-2B-it, Gemma-2-9B-it, Llama-3.1-8B-Instruct · paper: SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
structure: Linear Direction · models: DiffRhythm, EnCodec, WavTokenizer, Stable Audio Open · paper: Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
structure: Concept Crystals (Parallelogram/Trapezoid Structure) · models: Gemma-2-2B, Gemma-2-9B · paper: The Geometry of Concepts: Sparse Autoencoder Feature Structure
structure: Linear Direction · models: Custom frozen multi-quadrotor navigation policy · paper: Steering Multirobot Behavior via Closed-Loop Affine Activation Editing
structure: Feature Lobes (Spatial-Functional Modularity) · models: Gemma-2-2B · paper: The Geometry of Concepts: Sparse Autoencoder Feature Structure
structure: Linear Direction · models: Llama-3-8B, Aya-23-8B · paper: Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
structure: Linear Direction · models: CLIP ViT-B/16, DINOv2 ViT-B/14 · paper: Interpretable and Testable Vision Features via Sparse Autoencoders
structure: Linear Direction · models: SDXL Turbo, Stable Diffusion XL · paper: One-Step Is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
structure: Circle, Affine Subspace, Paraboloid (circular × pinched continuum) · models: Llama-3.1-8B · paper: Do Sparse Autoencoders Capture Concept Manifolds?
structure: Linear Direction · models: Stable Diffusion v1.4 · paper: Emergence and Evolution of Interpretable Concepts in Diffusion Models
structure: Linear Direction · models: MusicGen-Large, MusicGen-Small · paper: Discovering and Steering Interpretable Concepts in Large Generative Music Models
structure: Linear Subspace · models: AST (Audio Spectrogram Transformer, AudioSet-finetuned), HuBERT-base, WavLM-base-plus, MERT-v1-95M · paper: Sparse Autoencoders Make Audio Foundation Models More Explainable
structure: Linear Direction · models: pi0.5 (PaliGemma VLA backbone), OpenVLA-7B · paper: Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models
structure: Circle, Cone, Platonic Representation Hypothesis · models: GPT-2-small, Mistral-7B, Llama-3-8B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Llama-3.1-8B · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, Not All Language Model Features Are One-Dimensionally Linear, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Do Sparse Autoencoders Capture Concept Manifolds?
structure: Linear Separability · models: Llama-3.2-1B, Gemma-2-2B · paper: Multilingual Language Models Encode Script Over Linguistic Structure
structure: Linear Direction · models: GPT-4o, o3-mini · paper: Persona Features Control Emergent Misalignment
structure: Linear Separability · models: Qwen2.5-7B-Instruct, Llama-3.1-8B, Falcon3-7B · paper: Monitoring Emergent Reward Hacking During Generation via Internal Activations
structure: Linear Direction · models: Self-supervised ViT-S MAE (30.1M params, trained on 3M Euclid Q1 galaxy images, 90% masking) · paper: Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders
structure: Linear Direction · models: FLUX.1, Stable Diffusion 3.5 · paper: Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Gemma-2-9B-it · paper: Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models