MATH · IN · MODELS

Mlp probing

Techniquebeginner

A small multi-layer perceptron (typically 1–2 hidden layers with a nonlinearity) trained on frozen activations to predict a target property; used in place of a linear probe when the property may be decodable only nonlinearly, at the cost of the interpretability a single weight direction provides.

Used in (13 observations)

structure: Linear Separability · models: GPT-OSS-20B · paper: A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
structure: Linear Subspace · models: DINOv2 ViT-B/14, DINOv2-B (ViT-Base, 4 register tokens), CLIP ViT-L/14, MAE ViT-Base (Masked Autoencoder), iBOT ViT-B/16, Stable Diffusion v2.1 · paper: Probing the 3D Awareness of Visual Foundation Models
structure: Affine Subspace · models: DeBERTa-v2-xxlarge, GPT-Neo-1.3B · paper: More than Correlation: Do Large Language Models Learn Causal Representations of Space?
structure: Linear Separability · models: BERT-base-uncased, BERT-large-uncased, GPT-2-small, GPT-2-Large, GPT-2-XL, Pythia-6.9B, OLMo 2 7B, Gemma-2-2B, Qwen2.5-1.5B-Instruct, Llama-3.1-8B · paper: Finding Lexical Identity and Inflectional Morphology in Modern Language Models
structure: Linear Subspace · models: DINO ViT-B/16, MAE ViT-Base (Masked Autoencoder), ResNet-50 (SimCLR contrastive pretraining, ImageNet), ResNet-50 (MoCo v2, unsupervised contrastive pretraining, ImageNet) · paper: Probing the Mid-level Vision Capabilities of Self-Supervised Learning
structure: Linear Direction · models: OthelloGPT · paper: Emergent Linear Representations in World Models of Self-Supervised Sequence Models
structure: Linear Separability · models: DeepSeek-R1-Distill-Llama-8B, DeepSeek-R1-Distill-Llama-70B, DeepSeek-R1-Distill-Qwen-1.5B, DeepSeek-R1-Distill-Qwen-7B, DeepSeek-R1-Distill-Qwen-32B, QwQ-32B, Llama-3.1-8B-Instruct · paper: Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
structure: Linear Direction · models: Counter-Language Transformer (1 layer, 4 heads, causal-masked encoder) · paper: Emergent Stack Representations in Modeling Counter Languages Using Transformers
structure: Linear Direction · models: BGE-large-en-v1.5, all-mpnet-base-v2, all-MiniLM-L6-v2, Qwen3-Embedding-0.6B, Qwen2.5-3B-Instruct · paper: Probing Spectrum-Like Organization of States of Mind in Transformer Representation Spaces
structure: Linear Subspace · models: TabPFN v2 (tabular in-context-learning foundation model) · paper: TabPFN Through The Looking Glass: An Interpretability Study of TabPFN and Its Internal Representations
structure: Linear Separability · models: Jukebox, MusicGen-Small, MusicGen-Large · paper: Do Music Generation Models Encode Music Theory? (SynTheory)
structure: Linear Separability · models: Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B · paper: Decoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion