Trains a linear classifier or regressor on frozen activations to test whether a concept is linearly decodable — the main operational test for the Linear Representation Hypothesis.
Used in (141 observations)
structure: Linear Subspace · models: Llama-3.1-8B-Instruct, Qwen2-7B-Instruct · paper: Death by a Thousand Directions: Exploring the Geometry of Harmfulness in LLMs through Subconcept Probing
structure: Linear Direction, Linear Separability · models: Llama 3.3 70B Instruct, Qwen2.5-72B · paper: Understanding Moral Reasoning Trajectories in LLMs: Toward Probing-Based Explainability
structure: Linear Separability · models: Eternis-Forecaster-8B (RLVR post-trained from Qwen3-8B), GLM-4.5-Air, Qwen3-8B · paper: What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
structure: Linear Direction · models: Bielik-11B, Llama-PLLuM-12B, Gemma 4 12B, Mistral NeMo 12B (base), Llama-3.1-8B, Qwen3-14B · paper: Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering
structure: Linear Separability · models: Qwen3-32B · paper: Tool-Call Dependency Structure Is Linearly Decodable in LLM Agent Residual Streams
structure: Linear Direction, Linear Separability · models: Llama 3.3 70B Instruct · paper: Probing and Steering Evaluation Awareness of Language Models
structure: Linear Direction · models: OpenVLA-7B, pi0.5 (PaliGemma VLA backbone), DINOv2 ViT-B/14, CLIP ViT-B/32 · paper: What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies
structure: Linear Direction · models: Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Gemma-2-2B, Llama-3.2-3B, SmolLM3-3B, Qwen3-0.6B · paper: A Circuit for Predicting Hierarchical Structure In-Context in Large Language Models
structure: 1D continuum manifold, Linear Direction · models: Gemma-2-9B, Mistral-7B, Llama-3-70B-Instruct · paper: Latent Structure of Affective Representations in Large Language Models
structure: Linear Separability · models: Inception V3 (ImageNet-trained), ResNet-50 (supervised, ImageNet) · paper: Understanding Intermediate Layers Using Linear Classifier Probes
structure: Linear Direction · models: VideoMAE-base · paper: Causal Physics Steering in Video World Models via Concept Activation Vectors
structure: Linear Direction · models: AlphaZero (chess, ResNet policy/value network, self-play) · paper: Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
structure: Linear Direction · models: V-JEPA 2.1, V-JEPA 2, VideoPrism, VideoMAE-v2 · paper: Latent Video Prediction Learns Better World Models
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Llama-3.1-8B, Llama-3.1-70B-Instruct, Llama-3.2-1B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-3B-Instruct, Gemma-2-9B-it · paper: A Geometric Account of Activation Steering through Angle-Norm Decomposition
structure: Linear Separability · models: GPT-OSS-20B · paper: A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
structure: Linear Direction · models: Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, Vicuna-7B-v1.5 · paper: Linear Representations of Political Perspective Emerge in Large Language Models
structure: Linear Direction · models: GloVe (Wikipedia + Gigaword, uncased), word2vec (Google News, 300d) · paper: World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddings
structure: Linear Subspace, Linear Direction · models: Bayesian Wind Tunnel Transformer (bijection task, 6 layers, 6 heads, d_model=192), Bayesian Wind Tunnel Transformer (HMM filtering task, 9 layers, 8 heads, d_model=256) · paper: The Bayesian Geometry of Transformer Attention
structure: Conceptual Belief Space Hypothesis, Linear Subspace · models: Llama-3.1-8B-Instruct · paper: Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
structure: Linear Subspace · models: BERT-base-cased · paper: Visualizing and Measuring the Geometry of BERT
structure: Linear Direction, Linear Separability · models: ESM-2 (650M), ESM-2 (15B), ESM3 (1.4B, OPEN), ESM3 (98B, LARGE), ESMC (600M), ESMC (6B), ProGen2-base, EvoDiff OA-DM · paper: Viral Proteins Reveal Geometry of Protein Language Models
structure: Circle, Constructive Interference Hypothesis · models: BOWS Autoencoder (tied-weight, ReLU), BOWS Autoencoder (tied-weight, linear — no ReLU, baseline), BOWS Toy Transformer (1 block, 8 heads, d_model=768) · paper: From Data Statistics to Feature Geometry: How Correlations Shape Superposition
structure: Polytope (Simplex) · models: Gemma-2B, Llama-3-8B, Qwen3-4B, Mistral-7B-v0.3 · paper: The Geometry of Categorical and Hierarchical Concepts in Large Language Models, When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs
structure: Linear Separability · models: Llama-3-8B, Mistral-7B-v0.3 · paper: Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
structure: Linear Direction · models: Custom Transformer-VAE, Autoregressive MultiSlotting variant (trained from scratch on SELFIES-tokenized molecules) · paper: Molecules Meet Language: Confound-Aware Representation Learning and Chemical Property Steering in Transformer-VAE Latent Spaces
structure: Linear Direction · models: Audio-MAE, EnCodec · paper: Probing Spatial Structure in Pretrained Audio Representations
structure: Linear Direction · models: ChessGPT-8L-25M, ChessGPT-16L-50M · paper: Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
structure: Linear Separability · models: CLIP ViT-B/32, CLIP ViT-B/16, CLIP ViT-L/14 · paper: CLIP Behaves Like a Bag-of-Words Model Cross-modally but Not Uni-modally
structure: Concept lattice (Formal Concept Analysis half-space model), Cone · models: Llama-3.1-8B, Gemma-7B, Mistral-7B, LLaMA-3 3B, Llama-3-8B, Llama-3-70B · paper: The Lattice Representation Hypothesis of Large Language Models
structure: Linear Subspace, Linear Direction · models: GPT-2 Small, GPT-2 Medium, GPT-2 Large, Qwen2-1.5B-Instruct, Qwen2-7B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Gemma-2-2B-it · paper: The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
structure: Linear Subspace · models: CosyVoice2 · paper: A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
structure: Linear Direction · models: Qwen2-VL-7B-Instruct, Gemma 3 4B Instruct · paper: The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
structure: Linear Separability · models: Nucleotide Transformer (500M) · paper: Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.2-3B, Llama-3.2-1B, Qwen3-8B · paper: A Shared Geometry of Difficulty in Multilingual Language Models
structure: Linear Separability · models: Wan2.1-1.3B, CogVideoX-2B, LTX-Video (diffusion-based video generator), V-JEPA (video joint-embedding predictive architecture), VideoMAE-base · paper: The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show
structure: Linear Direction · models: DRC(3,3) Sokoban Agent · paper: Interpreting Emergent Planning in Model-Free Reinforcement Learning
structure: Linear Direction · models: Llama 3.3 70B Instruct, Qwen3-8B, Qwen2.5-14B-Instruct · paper: When Roleplaying, Do Models Believe What They Say?
structure: Linear Subspace · models: Custom 3-layer decision-pretraining Transformer (Fang & Rajan 2026) · paper: From Memories to Maps: Mechanisms of In-Context Reinforcement Learning in Transformers
structure: Linear Separability, Linear Subspace · models: Whisper base, Wav2Vec 2.0 Base (LibriSpeech-960h) · paper: Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages
structure: Linear Direction · models: Llama 3.3 70B Instruct · paper: Linear Personality Probing and Steering in LLMs: A Big Five Study
structure: Linear Subspace · models: GPT-2-small, GPT-OSS-20B, Llama-3-8B, Llama 4 Scout (17B-active/16-expert MoE), DeepSeek-V3, Mamba-2.8B, Falcon-Mamba-7B, xLSTM-7B, Kimi Linear 48B-A3B, GloVe (Wikipedia + Gigaword, uncased), FastText (bag-of-word-vectors), Convergent-Evolution 300M controlled Transformer, Convergent-Evolution 300M controlled LSTM, Convergent-Evolution 300M controlled Linear RNN (Gated DeltaNet) · paper: Convergent Evolution: How Different Language Models Learn Similar Number Representations
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-9B, Llama-3.1-8B, PaliGemma2 3B mix-448, PaliGemma2 10B mix-448, Idefics3-8B-Llama3 · paper: Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
structure: Curvature profile of the representation manifold, Linear Subspace · models: Llama-3.2-1B · paper: The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
structure: Linear Direction · models: Gemma-2-2B, Gemma-2-2B-it · paper: A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
structure: Linear Direction · models: Gemma 3 4B Instruct, Llama-3.2-3B-Instruct · paper: Sycophancy Hides Linearly in the Attention Heads
structure: Affine Subspace · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Pythia-6.9B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia) · paper: Language Models Represent Space and Time, Symmetry in Language Statistics Shapes the Geometry of Model Representations
structure: Affine Subspace · models: DeBERTa-v2-xxlarge, GPT-Neo-1.3B · paper: More than Correlation: Do Large Language Models Learn Causal Representations of Space?
structure: Decision boundary (as a codimension-1 hypersurface), Linear Direction · models: Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, Gemma-2-9B-it, Qwen2.5-7B-Instruct, Phi-3.5-mini-instruct, Llama-3-8B · paper: Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-2-7B-Chat · paper: Probing then Editing Response Personality of Large Language Models
structure: Linear Direction · models: Llama-3-8B-Instruct, Pythia-70M · paper: Inference-Time Causal Probing in LLMs
structure: Linear Direction · models: PGGAN (trained on CelebA-HQ, 1024x1024), StyleGAN (trained on FFHQ, 1024x1024) · paper: InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs
structure: Linear Direction · models: OpenFlamingo-3B-Instruct, OpenFlamingo-4B, OpenFlamingo-9B · paper: Visual concept ranking uncovers medical shortcuts used by large multimodal models
structure: Circle, Linear Direction · models: V-JEPA 2, VideoMAE-v2 · paper: Interpreting Physics in Video World Models
structure: Belief State Geometry Hypothesis (Mixed-State Presentation) · models: Custom GPT-2-style transformer (87M params, trained on synthetic Poker Hand History trajectories) · paper: Emergent World Beliefs: Exploring Transformers in Stochastic Games
structure: Linear Separability · models: Llama-3.1-8B, Qwen-2.5-7B, Qwen2.5-Math-7B, OpenMath2-Llama3.1-8B, OpenR1-Qwen-7B, GPT-OSS-20B · paper: How Language Directions Align with Token Geometry in Multilingual LLMs
structure: Linear Direction, Linear Separability · models: Gaperon-8B · paper: Language-Switching Triggers Take a Latent Detour Through Language Models
structure: Linear Separability · models: Qwen2.5-VL 7B Instruct, Qwen2.5-VL 32B Instruct, InternVL3-8B · paper: Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
structure: Linear Direction · models: T5-11B, UnifiedQA-11B (T5-based), T0++, GPT-J-6B, RoBERTa-large-MNLI, DeBERTa-xxlarge-v2-MNLI · paper: Discovering Latent Knowledge in Language Models Without Supervision
structure: Linear Separability · models: V-JEPA (video joint-embedding predictive architecture), VideoMAE-base, LTX-Video (diffusion-based video generator) · paper: Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis
structure: Linear Separability, Linear Direction · models: Llama-3-8B, Llama 3.1 70B · paper: Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
structure: Linear Direction, Linear Subspace · models: BART-base, T5-Base · paper: Implicit Representations of Meaning in Neural Language Models
structure: Linear Separability · models: BERT-base-uncased, BERT-large-uncased, GPT-2-small, GPT-2-Large, GPT-2-XL, Pythia-6.9B, OLMo 2 7B, Gemma-2-2B, Qwen2.5-1.5B-Instruct, Llama-3.1-8B · paper: Finding Lexical Identity and Inflectional Morphology in Modern Language Models
structure: Linear Direction, Linear Separability · models: Qwen3-0.6B, Qwen3-4B, Qwen3-14B, DeepSeek-R1-Distill-Qwen-7B, Llama-3.1-8B, Llama-2-7B, OLMo 3 7B, Gemma 3 4B Instruct, GPT-4o, GPT-OSS-20B · paper: What Really Controls Temporal Reasoning in LLMs: Tokenisation or Representation of Time?
structure: Linear Separability · models: ChemBERTa-2 (77M molecules), MolFormer (100M molecules), RoBERTa-Zinc-480M · paper: Probing Chemical Language Models: Effects of Pre-training and Fine-tuning
structure: Linear Separability · models: Qwen3-1.7B, Qwen3-4B-Instruct, Llama-3.1-8B-Instruct, Qwen3-14B, Qwen3-32B, Llama 3.3 70B Instruct · paper: LLM Agents Already Know When to Call Tools — Even Without Reasoning
structure: Linear Separability, Linear Direction · models: Llama-2-7B-Chat, Vicuna-7B · paper: Unlocking the Future: Look-Ahead Planning Mechanistic Interpretability in LLMs
structure: Linear Subspace · models: GRU passive object-state world model (Liu & Chen 2026), RSSM-lite passive object-state world model (Liu & Chen 2026) · paper: Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models
structure: Linear Subspace · models: Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct · paper: From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task
structure: Linear Subspace, Linear Direction · models: CodeLlama-13B, Gemma-2-2B, Llama 3.1 70B · paper: Do Language Models Track Entities Across State Changes?
structure: Linear Subspace · models: BERT-large-uncased, RoBERTa-large, ELECTRA-large (discriminator), BERT-mini, BERT-small, BERT-medium, BERT-base-uncased · paper: Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
structure: 1D continuum manifold · models: Pythia-2.8B, Llama-2-7B, Llama-3.1-8B, Llama-3.2-1B, GPT-2-Large, Mistral-7B, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Gemma-2-2B, Llama-3.1-70B-Instruct, Llama-3-8B-Instruct, Llama-3-8B, Mistral-7B-Instruct-v0.3, Qwen2.5-7B-Instruct · paper: Number Representations in LLMs: A Computational Parallel to Human Perception, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models, LLMs Know More About Numbers than They Can Say
structure: Linear Separability, Anisotropy · models: Llama-3-8B, Llama-2-7B · paper: Analysing the Residual Stream of Language Models Under Knowledge Conflicts
structure: Linear Representation Hypothesis · models: Llama-2-7B, Gemma-2B · paper: The Linear Representation Hypothesis and the Geometry of Large Language Models
structure: Laguerre-Voronoi partition (weighted power diagram of a linear readout layer) · models: Phi-2, Gemma-2-2B, Gemma-2-9B, Gemma-3-270M-it, Pythia-70M, Llama-3.1-8B · paper: Laguerre Geometry for Interpreting Large Language Models
structure: Linear Direction · models: Llama 3.3 70B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-7B-Instruct · paper: The Truthfulness Spectrum Hypothesis
structure: 1D continuum manifold, Linear Subspace · models: Llama-2-7B · paper: Probing for Representation Manifolds in Superposition
structure: Linear Direction · models: LAION-CLAP (HTSAT backbone, trained on LAION-Audio-630k + AudioSet + speech/music) · paper: Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
structure: Linear Direction · models: hallway (maze-solving transformer, forkless mazes), jirpy (maze-solving transformer, forking mazes) · paper: Structured World Representations in Maze-Solving Transformers
structure: Linear Separability · models: DINO ViT-S/16 trained on SSL4EO (Sentinel-1/2 satellite imagery) · paper: Probing Geospatial SSL Representations with Environmental Signals
structure: Linear Direction · models: Llama-2-7B, Llama-2-13B, Llama-3.1-8B, Llama-3.1-70B-Instruct, Mistral-7B-v0.1 · paper: Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
structure: Linear Subspace, Linear Direction · models: BERT-base-uncased, GPT-2 Small, Qwen-2.5-7B, Qwen2.5-Math-7B · paper: The Representational Geometry of Number
structure: Linear Direction · models: OpenVLA-7B · paper: Emergent World Representations in OpenVLA
structure: Linear Direction · models: OthelloGPT · paper: Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT
structure: Linear Direction · models: OthelloGPT · paper: Emergent Linear Representations in World Models of Self-Supervised Sequence Models
structure: Linear Direction · models: DiffRhythm, EnCodec, WavTokenizer, Stable Audio Open · paper: Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
structure: Linear Direction · models: Chronos, MOMENT · paper: On the Internal Semantics of Time-Series Foundation Models
structure: Linear Direction · models: Gemma 3 4B, MetaCLIP 2 · paper: The Information Geometry of Softmax: Probing and Steering
structure: Linear Separability · models: AION-1 (omnimodal foundation model for astronomical sciences) · paper: AION-1: Omnimodal Foundation Model for Astronomical Sciences
structure: Linear Subspace · models: Custom GPT-2-Small (trained from scratch on synthetic grid-navigation token sequences) · paper: Cognitive Maps in Language Models: A Mechanistic Analysis of Spatial Planning
structure: Linear Separability · models: GCN / GAT / GIN (trained on synthetic Grid-House and ClinTox molecular graphs) · paper: Do Graph Neural Network States Contain Graph Properties?
structure: Linear Subspace · models: Llama-3-8B-Instruct, Llama-3.1-8B-Instruct, Aya Expanse 8B, Gemma 2 9B IT, Qwen2.5-7B-Instruct · paper: Understanding Subword Compositionality of Large Language Models
structure: Linear Direction, Polytope (Simplex) · models: Pythia-160M · paper: (How) Do Language Models Track State?
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-2-7B, Llama-2-7B-Chat, Vicuna-7B, GPT-J-6B, ChatGLM3-6B, Qwen-2.5-7B, Qwen2.5-7B-Instruct, InternLM-7B, InternLM2-7B · paper: Probing then Editing Response Personality of Large Language Models
structure: Linear Separability · models: Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, Qwen3-4B-Instruct · paper: Tracing Relational Knowledge Recall in Large Language Models
structure: Linear Direction · models: Byte mLSTM (4096-unit, trained on ~82M Amazon reviews) · paper: Learning to Generate Reviews and Discovering Sentiment
structure: Linear Subspace · models: Qwen2.5-0.5B, Qwen3-14B, Llama-3.1-8B, Llama-3.2-1B, Phi-4 (15B) · paper: Geometric Factual Recall in Transformers
structure: Linear Separability · models: DeepSeek-R1-Distill-Llama-8B, DeepSeek-R1-Distill-Llama-70B, DeepSeek-R1-Distill-Qwen-1.5B, DeepSeek-R1-Distill-Qwen-7B, DeepSeek-R1-Distill-Qwen-32B, QwQ-32B, Llama-3.1-8B-Instruct · paper: Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
structure: Linear Subspace · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, OLMo 2 7B, OLMo 2 7B Instruct, Ministral-8B-Instruct-2410 · paper: Emotions Where Art Thou: Characterizing the Emotional Latent Space of LLMs
structure: Linear Separability · models: Custom RES-L-H ResNet (trained from scratch with VICReg/SimCLR objectives on CIFAR-100/FOOD101) · paper: Reverse Engineering Self-Supervised Learning
structure: 1D continuum manifold · models: Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Claude 3.5 Haiku · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, When Models Manipulate Manifolds: The Geometry of a Counting Task
structure: Linear Direction · models: ViT-B/16 (ImageNet-21k, supervised) · paper: From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers
structure: Linear Separability · models: Dolph2Vec (Wav2Vec2.0 architecture adapted for 44.1kHz dolphin audio, trained on 100hrs/5yr real longitudinal recordings) · paper: Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations
structure: Linear Direction · models: GPT-2 Small, Pythia-1.4B, Pythia-2.8B · paper: Linear Representations of Sentiment in Large Language Models
structure: Linear Direction, Linear Separability · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Llama-3-8B, Llama-3-70B, Gemma-2B, Gemma-7B · paper: Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
structure: Linear Direction · models: Qwen3.6-35B-A3B (MoE, 35B total / 3B active), Laguna-XS.2 · paper: Latent Programming Horizons in Coding Agents
structure: 1D continuum manifold · models: 1-Layer Ordinal Local-Comparison Transformer, Qwen2.5-1.5B · paper: Emergent Ordinal Geometry in Transformers Trained on Local Comparisons
structure: Linear Direction · models: DRC(3,3) Sokoban Agent · paper: Planning in a Recurrent Neural Network That Plays Sokoban
structure: Linear Direction, Concept Crystals (Parallelogram/Trapezoid Structure) · models: Llama-3.2-3B-Instruct, Llama-3.2-1B-Instruct, Qwen3-1.7B · paper: Linear Spatial World Models Emerge in Large Language Models
structure: Linear Subspace · models: HuBERT Large (LibriLight 60k), wav2vec 2.0 Large (LV-60), XLS-R (300M, 53 languages), MMS (1B, 1000+ languages) · paper: Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
structure: Linear Direction · models: Counter-Language Transformer (1 layer, 4 heads, causal-masked encoder) · paper: Emergent Stack Representations in Modeling Counter Languages Using Transformers
structure: Linear Direction · models: BGE-large-en-v1.5, all-mpnet-base-v2, all-MiniLM-L6-v2, Qwen3-Embedding-0.6B, Qwen2.5-3B-Instruct · paper: Probing Spectrum-Like Organization of States of Mind in Transformer Representation Spaces
structure: Linear Subspace, Linear Direction · models: Gemma-2-2B-it, Llama-2-7B-Chat · paper: The Cylindrical Representation Hypothesis for Language Model Steering
structure: Linear Separability · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, DeepSeek-R1-Distill-Llama-8B · paper: LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
structure: Linear Subspace · models: DINOv2 ViT-L/14, MAE ViT-Large (Masked Autoencoder), iBOT ViT-L/16 · paper: Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
structure: Linear Separability · models: OpenVLA-7B · paper: Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
structure: Linear Subspace · models: TabPFN v2 (tabular in-context-learning foundation model) · paper: TabPFN Through The Looking Glass: An Interpretability Study of TabPFN and Its Internal Representations
structure: Linear Direction · models: Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, Phi-4-Mini-Instruct · paper: The Geometries of Truth Are Orthogonal Across Tasks
structure: Linear Direction · models: GoogLeNet (ImageNet-trained), Inception V3 (ImageNet-trained) · paper: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
structure: Linear Separability · models: Silverman 4x5 CNN (self-play chess RL agent), Los Alamos 6x6 ResNet-CNN (self-play chess RL agent) · paper: Reinforcement Learning in an Adaptable Chess Environment for Detecting Human-understandable Concepts
structure: Linear Direction · models: Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-2-9B-it, Gemma-2-2B-it · paper: The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
structure: Linear Direction · models: Qwen3-4B-Instruct-2507 · paper: Temporal Preference Concepts and Their Functions in a Large Language Model
structure: Intrinsic-dimension profile across depth, Linear Separability · models: Qwen2.5-0.5B-Instruct, Qwen2.5-1.5B, Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-32B, Gemma-2-2B, Gemma-2-9B, Gemma-2-27B, Llama-3.2-1B, Llama-3.2-3B · paper: Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs
structure: Linear Separability, Linear Direction · models: Mistral Small 3 (2501), Mistral-7B, Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-3.2-3B, Gemma-2-9B, Gemma-2-2B, GPT-J-6B, GPT-2 XL, GPT-2 Large, GPT-2 Medium, GPT-2 Small · paper: Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
structure: Circle, Cone, Platonic Representation Hypothesis · models: GPT-2-small, Mistral-7B, Llama-3-8B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Llama-3.1-8B · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, Not All Language Model Features Are One-Dimensionally Linear, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Do Sparse Autoencoders Capture Concept Manifolds?
structure: Linear Direction · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-v0.1 · paper: The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets, On the Universal Truthfulness Hyperplane Inside LLMs
structure: Linear Direction · models: LLaMA-7B, Alpaca-7B, Vicuna-7B · paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
structure: Linear Direction · models: ESM-2 (650M) · paper: Sparse Autoencoders for Low-N Protein Function Prediction and Design
structure: Circle, Linear Direction · models: Llama-3.1-8B-Instruct, Qwen3-8B, Qwen3-14B, Apertus-8B-Instruct-2509, Gemma 4 E4B-it · paper: Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control, Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
structure: Linear Separability · models: TapeBert (BERT-Base, Pfam-pretrained), ProtBert, ProtBert-BFD, ProtAlbert, ProtXLNet · paper: BERTology Meets Biology: Interpreting Attention in Protein Language Models
structure: Anisotropy · models: Mamba-130m (state-spaces/mamba-130m-hf), RoBERTa-base · paper: Lost in State Space: Probing Frozen Mamba Representations
structure: Linear Separability · models: Gemma-2-2B-it, Qwen2.5-3B, Qwen2.5-7B, Qwen2.5-14B, Llama-3.1-8B · paper: Two Axes of LLM Abstention: Answer Correctness and Question Answerability
structure: Linear Centroids Hypothesis · models: ResNet-50 (supervised, ImageNet), DINOv2 ViT-L/14, DINOv3 ViT-B/16, GPT-2-Large, Llama-3.1-8B · paper: The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts
structure: Linear Separability · models: DAC (Descript Audio Codec), SpeechTokenizer · paper: Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
structure: Linear Separability · models: Qwen2.5-7B-Instruct, Llama-3.1-8B, Falcon3-7B · paper: Monitoring Emergent Reward Hacking During Generation via Internal Activations
structure: Linear Direction · models: Chronos, MOMENT, Moirai-1.1-R Large · paper: Exploring Representations and Interventions in Time Series Foundation Models
structure: Linear Subspace, Attention–MLP Sufficiency Staging Hypothesis · models: Grid-walker decoder transformer (L4/H4/d_model=128, HookedTransformer) · paper: Predictive Statistics Shape Emergent World Representations of Grid Walkers
structure: Linear Direction, Constraint-Algebra Basis Hypothesis · models: Sudoku Transformer (8 layers, 8 heads, d_model=576) · paper: Transformers Linearly Represent Highly Structured World Models
structure: Linear Direction · models: Qwen3-32B, Llama 3.3 70B Instruct · paper: Rhetorical Questions in LLM Representations: A Linear Probing Study
structure: Linear Direction · models: Qwen3.5-35B-A3B (MoE, ~3B active/token) · paper: Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits
structure: Linear Direction · models: IRIS (Atari 100k, Breakout/Pong), DIAMOND (Atari 100k, Breakout/Pong) · paper: What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
structure: Linear Separability, Linear Direction · models: Mistral-7B-Instruct, DeepSeek-LLM-7B-Chat · paper: Language Models Represent Beliefs of Self and Others