Pythia
EleutherAI
Structures found in this family (17)
Linear Subspace, Linear Direction, Intrinsic-dimension profile across depth, Dimensional collapse, Cone, Tree Metric Embedding, Persistent-homology / Betti profile across depth, Affine Subspace, Generalized Helix, Platonic Representation Hypothesis, Linear Separability, 1D continuum manifold, Laguerre-Voronoi partition (weighted power diagram of a linear readout layer), Polytope (Simplex), Anisotropy, Tangent-Aligned Anisotropy Hypothesis, Attention reference frame (sink-token anchor configuration)
Hypotheses argued for by this family (2)
By model (10)
Pythia-160M · 160M
Pythia-6.9B · 6.9B
Pythia-2.8B · 2.8B
Pythia-410M · 410M
Pythia-1.4B · 1.4B
Pythia-70M · 70M
Pythia-1B · 1B
Pythia-12B · 12B
Pythia-160M-deduped · 160M
Pythia-70M-deduped · 70M
Observations (47)
Papers
Representational Analysis of Binding in Language Models (2024), Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability (2025), Functional Subspace, where language models can use vector algebra to solve problems (2026), Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning (2025), Investigating Representation Universality: Case Study on Genealogical Representations (2024), Transferring Linear Features Across Language Models With Model Stitching (2025), Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs (2024), Language Models Represent and Transform Concepts with Shared Geometry (2026), Polar probe linearly decodes semantic structures from LLMs (2026), Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models (2025), How Do Language Models Bind Entities in Context? (2023), Function Vectors in Large Language Models (2024), In-Context Learning Creates Task Vectors (2023), Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning (2025), Do Different Prompting Methods Yield a Common Task Representation in Language Models? (2025), Persistent Topological Features in Large Language Models (2024), Improving Dictionary Learning with Gated Sparse Autoencoders (2024), Language Models Represent Space and Time (2024), Symmetry in Language Statistics Shapes the Geometry of Model Representations (2026), Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence (2025), Inference-Time Causal Probing in LLMs (2026), Language Models Use Trigonometry to Do Addition (2025), Representational Curvature Modulates Behavioral Uncertainty in Large Language Models (2026), Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders (2024), Semantic Convergence: Investigating Shared Representations Across Scaled LLMs (2025), Abstraction Induces the Brain Alignment of Language and Speech Models (2026), Geometric Signatures of Compositionality Across a Language Model's Lifetime (2024), Finding Lexical Identity and Inflectional Morphology in Modern Language Models (2025), Rethinking Intrinsic Dimension Estimation in Neural Representations (2026), Number Representations in LLMs: A Computational Parallel to Human Perception (2025), Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling (2025), Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models (2026), LLMs Know More About Numbers than They Can Say (2026), Laguerre Geometry for Interpreting Large Language Models (2026), Attention Sinks and Compression Valleys in LLMs Are Two Sides of the Same Coin (2025), How Reliable are Causal Probing Interventions? (2025), LEACE: Perfect Linear Concept Erasure in Closed Form (2023), Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026), (How) Do Language Models Track State? (2025), Outlier Dimensions Encode Task-Specific Knowledge (2023), Tracing the Representation Geometry of Language Models from Pretraining to Post-training (2025), The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models (2024), Linear Representations of Sentiment in Large Language Models (2023), Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks (2026), Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics (2026), Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings (2026), Transcoders Find Interpretable LLM Feature Circuits (2024), What are you sinking? A geometric approach on attention sink (2025), Eliciting Latent Predictions from Transformers with the Tuned Lens (2023), Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning (2026), Scale Determines Whether Language Models Organize Representation Geometry for Prediction (2026), Which Attention Heads Matter for In-Context Learning? (2025), When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception (2026)