MATH · IN · MODELS

Gemma

Google DeepMind

By model (29)

Gemma 2 27B Instruct · 27B
Gemma 4 E4B-it · 4B
Gemma 2 9B IT · 9B
Gemma 1.1 7B Instruct · 7B
Gemma 3 12B Instruct · 12B
Gemma-3-1B-it · 1B
Gemma 4 12B · 12B
Gemma-4-31B-Instruct · 31B
Gemma-7B-it · 7B
Gemma 3 12B IT · 12B
Gemma 3 1B IT · 1B
Gemma-2B-it · 2B
no structures recorded for this checkpoint specifically
Gemma 2 9B · 9B
no structures recorded for this checkpoint specifically

Observations (122)

Papers

Geometry of Ordinal Representations in Language Models (2026), When Models Manipulate Manifolds: The Geometry of a Counting Task (2025), Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal (2026), The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models (2026), Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering (2026), A Circuit for Predicting Hierarchical Structure In-Context in Large Language Models (2025), Tool Calling Is Linearly Readable and Steerable in Language Models (2026), Concept Heterogeneity-aware Representation Steering (2026), Latent Structure of Affective Representations in Large Language Models (2026), Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation (2026), A Minimalist Structural Probe for Phase-Count Abstraction in Transformer Language Models (2026), Functional Subspace, where language models can use vector algebra to solve problems (2026), A Geometric Account of Activation Steering through Angle-Norm Decomposition (2026), Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning (2025), Investigating Representation Universality: Case Study on Genealogical Representations (2024), Dissociating the Internal Representations of Sycophancy in LLMs (2026), Head Pursuit: Probing Attention Specialization in Multimodal Transformers (2025), BatchTopK Sparse Autoencoders (2024), Finding Belief Geometries with Sparse Autoencoders (2026), Conceptors for Semantic Steering (2026), The Geometry of Categorical and Hierarchical Concepts in Large Language Models (2024), When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs (2026), A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders (2024), Transferring Linear Features Across Language Models With Model Stitching (2025), Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations (2026), A Mechanistic Investigation of Supervised Fine Tuning (2026), Causal Language Control in Multilingual Transformers via Sparse Feature Steering (2025), Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations (2026), The Lattice Representation Hypothesis of Large Language Models (2026), The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models (2026), Context Structure Reshapes the Representational Geometry of Language Models (2026), Language Models Represent and Transform Concepts with Shared Geometry (2026), Unveiling the Latent Directions of Reflection in Large Language Models (2025), The Dual Mechanisms of Spatial Reasoning in Vision-Language Models (2026), Scenario-based Probing and Steering Cultural Values in Large Language Models (2026), Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings (2026), GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs (2026), Causal Probing for Internal Visual Representations in Multimodal Large Language Models (2026), How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning (2026), Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models (2025), Improving Dictionary Learning with Gated Sparse Autoencoders (2024), A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations (2026), Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders (2025), Sycophancy Hides Linearly in the Attention Heads (2026), Language Models Represent Space and Time (2024), Symmetry in Language Statistics Shapes the Geometry of Model Representations (2026), Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling (2025), The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws (2026), Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence (2025), Revisiting the Platonic Representation Hypothesis: An Aristotelian View (2026), How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? (2026), Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models (2026), Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries (2026), High-Dimensional Interlingual Representations of Large Language Models (2025), Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations (2026), In-Context Learning of Representations (2025), The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024), Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency (2026), Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding (2025), Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention (2026), Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders (2024), Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models (2026), Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes (2026), Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders (2024), Semantic Convergence: Investigating Shared Representations Across Scaled LLMs (2025), Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence (2026), Shared Global and Local Geometry of Language Model Embeddings (2025), Just-in-Time and Distributed Task Representations in Language Models (2025), Finding Lexical Identity and Inflectional Morphology in Modern Language Models (2025), What Really Controls Temporal Reasoning in LLMs: Tokenisation or Representation of Time? (2026), Catching Rationalization in the Act: Detecting Motivated Reasoning Before and After CoT via Activation Probing (2026), Do Language Models Track Entities Across State Changes? (2026), Number Representations in LLMs: A Computational Parallel to Human Perception (2025), Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models (2026), LLMs Know More About Numbers than They Can Say (2026), The Linear Representation Hypothesis and the Geometry of Large Language Models (2023), Relational Linearity is a Predictor of Hallucinations (2026), Universal Response and Emergence of Induction in LLMs (2024), Laguerre Geometry for Interpreting Large Language Models (2026), Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs (2026), Attention Sinks and Compression Valleys in LLMs Are Two Sides of the Same Coin (2025), Learning Multi-Level Features with Matryoshka Sparse Autoencoders (2025), Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning (2025), Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences (2026), Understanding Emergent Misalignment via Feature Superposition Geometry (2026), SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models (2025), Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens (2026), Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Activation Monitors (2025), Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy (2026), A Retrieval-Conditioned Rebinding Circuit for Dynamic Entity Tracking in Large Language Models (2026), The Information Geometry of Softmax: Probing and Steering (2026), Understanding Subword Compositionality of Large Language Models (2025), Atlas-Alignment: Making Interpretability Transferable Across Language Models (2025), Reasoning emerges from constrained inference manifolds in large language models (2026), The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence (2025), Refusal Direction is Universal Across Safety-Aligned Languages (2025), Refusal in Language Models Is Mediated by a Single Direction (2024), Representation Engineering: A Top-Down Approach to AI Transparency (2023), Programming Refusal with Conditional Activation Steering (2024), The Platonic Representation Hypothesis (2024), Beyond Position: the emergence of wavelet-like properties in Transformers (2025), Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 (2024), Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning (2025), Language Models Are Implicitly Continuous (2025), The Cylindrical Representation Hypothesis for Language Model Steering (2026), Steered LLM Activations Are Non-Surjective (2026), ReCoVeR the Target Language: Language Steering Without Sacrificing Task Performance (2025), Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families (2026), The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes (2026), The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations (2026), The Blessing and Curse of Dimensionality in Safety Alignment (2025), Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs (2026), Large Language Models Encode Semantics and Alignment in Linearly Separable Representations (2025), Not All Language Model Features Are One-Dimensionally Linear (2024), Do Sparse Autoencoders Capture Concept Manifolds? (2026), Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology (2026), What are you sinking? A geometric approach on attention sink (2025), From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs (2025), Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control (2026), Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs (2026), Multilingual Language Models Encode Script Over Linguistic Structure (2026), The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models (2026), Two Axes of LLM Abstention: Answer Correctness and Question Answerability (2026), The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models (2026), Scale Determines Whether Language Models Organize Representation Geometry for Prediction (2026), Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models (2025), The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models (2025), When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception (2026)