MATH · IN · MODELS

GPT

OpenAI

Observations (40)

Papers

Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits (2025), How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings (2019), Investigating Aspect Features in Contextualized Embeddings with Semantic Scales and Distributional Similarity (2024), Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models (2025), A Geometric Notion of Causal Probing (2023), Transferring Linear Features Across Language Models With Model Stitching (2025), Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs (2024), Isotropy in the Contextual Embedding Space: Clusters and Manifolds (2021), Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models (2025), Emergence of Separable Manifolds in Deep Language Representations (2020), Convergent Evolution: How Different Language Models Learn Similar Number Representations (2026), Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence (2025), Pre-trained Large Language Models Use Fourier Features to Compute Addition (2024), The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024), IsoScore: Measuring the Uniformity of Embedding Space Utilization (2021), Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding (2025), Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing (2024), Shared Global and Local Geometry of Language Model Embeddings (2025), Finding Lexical Identity and Inflectional Morphology in Modern Language Models (2025), What Really Controls Temporal Reasoning in LLMs: Tokenisation or Representation of Time? (2026), Number Representations in LLMs: A Computational Parallel to Human Perception (2025), Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling (2025), Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models (2026), LLMs Know More About Numbers than They Can Say (2026), Universal Response and Emergence of Induction in LLMs (2024), The Origins of Representation Manifolds in Large Language Models (2025), Copy Suppression: Comprehensively Understanding an Attention Head (2023), How Reliable are Causal Probing Interventions? (2025), Mapping Language Models to Grounded Conceptual Spaces (2022), Model Editing as a Robust and Denoised Variant of DPO: A Case Study on Toxicity (2024), The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models (2024), Intervention Lens: from Representation Surgery to String Counterfactuals (2024), All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality (2021), Multilinguality as Sense Adaptation (2026), Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings (2026), Scaling and Evaluating Sparse Autoencoders (2024), Symmetry in Language Statistics Shapes the Geometry of Model Representations (2026), Not All Language Model Features Are One-Dimensionally Linear (2024), Do Sparse Autoencoders Capture Concept Manifolds? (2026), The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts (2026), Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning (2026), Persona Features Control Emergent Misalignment (2026), Which Attention Heads Matter for In-Context Learning? (2025)