MATH · IN · MODELS

GPT-2

OpenAI

Hypotheses argued for by this family (1)

By model (8)

Custom GPT-2-Small (trained from scratch on synthetic grid-navigation token sequences) · 12-layer (GPT-2-Small architecture), 124M
nanoGPT-style GPT-2 124M (FineWeb-10B) · 124M
Custom 12-layer GPT-2-Small (trained from scratch, 100M tokens) · 12-layer (GPT-2-Small architecture)

Observations (28)

Papers

Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning (2025), BatchTopK Sparse Autoencoders (2024), Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models (2023), The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models (2026), Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space (2022), Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability (2026), Representational Curvature Modulates Behavioral Uncertainty in Large Language Models (2026), Language models align with brain regions that represent concepts across modalities (2025), Linearity of Relation Decoding in Transformer Language Models (2023), Invariant Reasoning Directions in Latent Trajectories of Language Models (2026), Language Models Implement Simple Word2Vec-style Vector Arithmetic (2024), The Representational Geometry of Number (2026), Trajectory Geometry of Transformer Representations Across Layers (2026), Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026), Mapping Language Models to Grounded Conceptual Spaces (2022), Cognitive Maps in Language Models: A Mechanistic Analysis of Spatial Planning (2025), Outlier Dimensions Encode Task-Specific Knowledge (2023), Locating and Editing Factual Associations in GPT (2022), Linear Representations of Sentiment in Large Language Models (2023), Language Models Are Implicitly Continuous (2025), The Blessing and Curse of Dimensionality in Safety Alignment (2025), Lines of Thought in Large Language Models (2024), Large Language Models Encode Semantics and Alignment in Linearly Separable Representations (2025), Scaling and Evaluating Sparse Autoencoders (2024), Transcoders Find Interpretable LLM Feature Circuits (2024), Revisiting the Othello World Model Hypothesis (2025), Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations (2025)