MATH · IN · MODELS
methods / Causal Validation / Logit Lens

Logit Lens

Techniquebeginner

Projects the residual stream after an intermediate layer directly through the model's final unembedding matrix, reading off what the model's 'current best guess' would be if generation stopped at that layer — a forward-pass-only diagnostic for tracking how a prediction builds up across depth.

Used in (11 observations)

structure: Linear Direction · models: Adaptive Input Transformer LM (16-layer, WikiText-103) · paper: Transformer Feed-Forward Layers Are Key-Value Memories
structure: Linear Subspace, Linear Direction · models: Leela Chess Zero (policy network) · paper: The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network
structure: Linear Direction · models: GPT-2 Small · paper: Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
structure: Lissajous Curves · models: GPT-2-XL · paper: Pre-trained Large Language Models Use Fourier Features to Compute Addition
structure: Linear Direction · models: OpenVLA-7B, pi0 (PaliGemma VLA backbone, predecessor of pi0.5) · paper: Mechanistic Interpretability for Steering Vision-Language-Action Models
structure: Linear Direction · models: Qwen3-1.7B, Llama-3.1-8B-Instruct, Gemma 3 1B IT, Qwen2.5-7B-Instruct · paper: Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
structure: Linear Direction · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Gemma-2-9B, Gemma-2-9B-it, Mistral-7B-v0.3, Mistral-7B-Instruct-v0.3 · paper: Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens
structure: Linear Direction · models: SmolLM-360M, OLMo-1B, OLMo-7B, Llama-2-7B · paper: Output Vector Editing for Memorization Mitigation in Large Language Models
structure: Linear Subspace · models: TabPFN v2 (tabular in-context-learning foundation model) · paper: TabPFN Through The Looking Glass: An Interpretability Study of TabPFN and Its Internal Representations
structure: Linear Direction, Constraint-Algebra Basis Hypothesis · models: Sudoku Transformer (8 layers, 8 heads, d_model=576) · paper: Transformers Linearly Represent Highly Structured World Models