MATH · IN · MODELS

Llama

Meta

By model (35)

GeLLM3O-LLaMA3
Llama-2-70B-Chat · 70B
Llama-3-8B-Lexi-Uncensored · 8B
Llama-3.1-Tulu-3-8B-DPO · 8B
Llama-3.1-Tulu-3-8B-SFT · 8B
Llama 4 Scout (17B-active/16-expert MoE) · 109B (17B active)
OpenMath2-Llama3.1-8B · 8B
Llama-3.2-11B-Instruct · 11B
no structures recorded for this checkpoint specifically
Llama-3.2-90B-Instruct · 90B
no structures recorded for this checkpoint specifically

Observations (210)

Papers

Semantic Structure of Feature Space in Large Language Models (2026), Death by a Thousand Directions: Exploring the Geometry of Harmfulness in LLMs through Subconcept Probing (2025), Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders (2024), Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal (2026), The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models (2026), The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models (2026), Understanding Moral Reasoning Trajectories in LLMs: Toward Probing-Based Explainability (2026), Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering (2026), Probing and Steering Evaluation Awareness of Language Models (2025), A Circuit for Predicting Hierarchical Structure In-Context in Large Language Models (2025), Cell-Based Representation of Relational Binding in Language Models (2026), On the Mutual Influence of Gender and Occupation in LLM Representations (2025), Representational Analysis of Binding in Language Models (2024), Tool Calling Is Linearly Readable and Steerable in Language Models (2026), Concept Heterogeneity-aware Representation Steering (2026), Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior (2026), Latent Structure of Affective Representations in Large Language Models (2026), Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation (2026), A Minimalist Structural Probe for Phase-Count Abstraction in Transformer Language Models (2026), The Geometry of Thought: How Scale Restructures Reasoning in Large Language Models (2026), Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs (2026), Functional Subspace, where language models can use vector algebra to solve problems (2026), A Geometric Account of Activation Steering through Angle-Norm Decomposition (2026), Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning (2026), Linear Representations of Political Perspective Emerge in Large Language Models (2025), Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning (2025), Investigating Representation Universality: Case Study on Genealogical Representations (2024), Dissociating the Internal Representations of Sycophancy in LLMs (2026), Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization (2026), Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space (2026), Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models (2026), Understanding (Un)Reliability of Steering Vectors in Language Models (2025), The Geometry of Categorical and Hierarchical Concepts in Large Language Models (2024), When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs (2026), A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders (2024), Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness (2026), Persona Vectors: Monitoring and Controlling Character Traits in Language Models (2025), Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations (2026), Do Sparse Autoencoders Capture Concept Manifolds? (2026), Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations (2026), The Lattice Representation Hypothesis of Large Language Models (2026), Representation Engineering: A Top-Down Approach to AI Transparency (2023), The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models (2026), Language Models Represent and Transform Concepts with Shared Geometry (2026), Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs (2025), Unravelling the Mechanisms of Manipulating Numbers in Language Models (2025), Scenario-based Probing and Steering Cultural Values in Large Language Models (2026), GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs (2026), Polar probe linearly decodes semantic structures from LLMs (2026), A Shared Geometry of Difficulty in Multilingual Language Models (2026), Label Words as Local Task Vectors in In-Context Learning (2024), Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models (2025), The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces (2024), When Roleplaying, Do Models Believe What They Say? (2026), How Do Language Models Bind Entities in Context? (2023), Arithmetic in the Wild: Llama Uses Base-10 Addition to Reason About Cyclic Concepts (2026), Linear Personality Probing and Steering in LLMs: A Big Five Study (2025), Convergent Evolution: How Different Language Models Learn Similar Number Representations (2026), How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning (2026), Function Vectors in Large Language Models (2024), In-Context Learning Creates Task Vectors (2023), Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning (2025), Do Different Prompting Methods Yield a Common Task Representation in Language Models? (2025), Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models (2025), The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors (2026), Sycophancy Hides Linearly in the Attention Heads (2026), Language Models Represent Space and Time (2024), Symmetry in Language Statistics Shapes the Geometry of Model Representations (2026), Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling (2025), Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence (2025), Characterizing Linear Alignment Across Language Models (2026), Revisiting the Platonic Representation Hypothesis: An Aristotelian View (2026), How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? (2026), Emergent Causal-Geometric Dynamics Across Depth in Large Language Models (2026), Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder-Decoder Perspective (2025), Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models (2026), Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries (2026), High-Dimensional Interlingual Representations of Large Language Models (2025), Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts (2025), The Effectiveness of Style Vectors for Steering LLMs: A Human Evaluation (2026), Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations (2026), Probing then Editing Response Personality of Large Language Models (2025), The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language Models (2024), In-Context Learning of Representations (2025), Inference-Time Causal Probing in LLMs (2026), The Geometry of Hidden Representations of Large Transformer Models (2023), The Geometry of Concepts: Sparse Autoencoder Feature Structure (2024), A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension (2024), Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency (2026), Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding (2025), Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models (2026), Language Models Use Trigonometry to Do Addition (2025), Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes (2026), How Language Directions Align with Token Geometry in Multilingual LLMs (2026), Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing (2024), Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders (2024), Semantic Convergence: Investigating Shared Representations Across Scaled LLMs (2025), Geometric Signatures of Compositionality Across a Language Model's Lifetime (2024), Shared Global and Local Geometry of Language Model Embeddings (2025), Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs (2025), Finding Lexical Identity and Inflectional Morphology in Modern Language Models (2025), What Really Controls Temporal Reasoning in LLMs: Tokenisation or Representation of Time? (2026), LLM Agents Already Know When to Call Tools — Even Without Reasoning (2026), Unlocking the Future: Look-Ahead Planning Mechanistic Interpretability in LLMs (2024), Catching Rationalization in the Act: Detecting Motivated Reasoning Before and After CoT via Activation Probing (2026), Rethinking Intrinsic Dimension Estimation in Neural Representations (2026), Language Models Encode Numbers Using Digit Representations in Base 10 (2025), From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task (2025), Revealing Emergent Human-like Conceptual Representations from Language Prediction (2025), Do Language Models Track Entities Across State Changes? (2026), Number Representations in LLMs: A Computational Parallel to Human Perception (2025), Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models (2026), LLMs Know More About Numbers than They Can Say (2026), Analysing the Residual Stream of Language Models Under Knowledge Conflicts (2024), Linearity of Relation Decoding in Transformer Language Models (2023), The Linear Representation Hypothesis and the Geometry of Large Language Models (2023), Relational Linearity is a Predictor of Hallucinations (2026), Universal Response and Emergence of Induction in LLMs (2024), Laguerre Geometry for Interpreting Large Language Models (2026), Do Linear Probes Generalize Better in Persona Coordinates? (2026), The Truthfulness Spectrum Hypothesis (2026), Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs (2026), Probing for Representation Manifolds in Superposition (2026), Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences (2026), Understanding Emergent Misalignment via Feature Superposition Geometry (2026), SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models (2025), Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens (2026), How Language Models Process Negation (2026), Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Activation Monitors (2025), How Reliable are Causal Probing Interventions? (2025), Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks (2025), LEACE: Perfect Linear Concept Erasure in Closed Form (2023), Steering Language Model Refusal with Sparse Autoencoder Features (2024), A Retrieval-Conditioned Rebinding Circuit for Dynamic Entity Tracking in Large Language Models (2026), Output Vector Editing for Memorization Mitigation in Large Language Models (2026), Understanding Subword Compositionality of Large Language Models (2025), Learning Uncertainty from Sequential Internal Dispersion in Large Language Models (2026), Steering at the Source: Style Modulation Heads for Robust Persona Control (2026), A polar coordinate system represents syntax in large language models (2024), Tracing Relational Knowledge Recall in Large Language Models (2026), Language Models Use Lookbacks to Track Beliefs (2025), Tracing the Representation Geometry of Language Models from Pretraining to Post-training (2025), Atlas-Alignment: Making Interpretability Transferable Across Language Models (2025), Simulated Adoption: Decoupling Magnitude and Direction in LLM In-Context Conflict Resolution (2026), Geometric Factual Recall in Transformers (2026), The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models (2024), The Geometry of Reasoning: Flowing Logics in Representation Space (2026), Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification (2025), Refusal in LLMs is an Affine Function (2024), The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence (2025), Refusal Direction is Universal Across Safety-Aligned Languages (2025), Refusal in Language Models Is Mediated by a Single Direction (2024), Programming Refusal with Conditional Activation Steering (2024), Emotions Where Art Thou: Characterizing the Emotional Latent Space of LLMs (2025), Relational Rank Geometry in Transformers: Detecting and Steering Hidden-State Relation Frames (2026), Identifying Linear Relational Concepts in Large Language Models (2023), The Platonic Representation Hypothesis (2024), Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction (2025), Beyond Position: the emergence of wavelet-like properties in Transformers (2025), Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages (2025), Understanding and Preserving Safety in Fine-Tuned LLMs (2026), Linear Representations of Hierarchical Concepts in Language Models (2026), Language Models Are Implicitly Continuous (2025), Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers (2025), Linear Spatial World Models Emerge in Large Language Models (2025), The Cylindrical Representation Hypothesis for Language Model Steering (2026), On the Non-Identifiability of Steering Vectors in Large Language Models (2026), Steered LLM Activations Are Non-Surjective (2026), ReCoVeR the Target Language: Language Steering Without Sacrificing Task Performance (2025), Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks (2026), LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals (2026), The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks (2026), Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs (2025), Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention (2026), Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families (2026), Task Recognition and Task Learning Heads Align In-Context Hidden States with a Label-Unembedding Task Subspace (2026), The Geometries of Truth Are Orthogonal Across Tasks (2026), The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes (2026), The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations (2026), The Blessing and Curse of Dimensionality in Safety Alignment (2025), Revisiting Hallucination Detection with Effective Rank-based Uncertainty (2025), Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs (2026), Lines of Thought in Large Language Models (2024), Large Language Models Encode Semantics and Alignment in Linearly Separable Representations (2025), Not All Language Model Features Are One-Dimensionally Linear (2024), Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology (2026), What are you sinking? A geometric approach on attention sink (2025), Internal states before "wait" modulate reasoning patterns (2025), The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets (2024), On the Universal Truthfulness Hyperplane Inside LLMs (2024), How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs (2026), Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (2023), Eliciting Latent Predictions from Transformers with the Tuned Lens (2023), Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control (2026), Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs (2026), Negative Before Positive: Asymmetric Valence Processing in Large Language Models (2026), Multilingual Language Models Encode Script Over Linguistic Structure (2026), Two Axes of LLM Abstention: Answer Correctness and Question Answerability (2026), The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts (2026), Monitoring Emergent Reward Hacking During Generation via Internal Activations (2026), The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities (2024), When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals (2026), Linear Relational Decoding of Morphology in Language Models (2025), Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models (2025), Rhetorical Questions in LLM Representations: A Linear Probing Study (2026), Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension (2024), Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models (2026), Which Attention Heads Matter for In-Context Learning? (2025), Spherical Steering: Geometry-Aware Activation Rotation for Language Models (2026), Tracing Moral Foundations in Large Language Models (2026), Revisiting the Othello World Model Hypothesis (2025), The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models (2025), SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing (2026), Decoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion (2025), Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2025)