MATH · IN · MODELS

Mistral

Mistral AI

By model (20)

Ministral-8B-Instruct-2410 · 8B
Mistral-7B-Instruct-v0.1 · 7.24B
Mistral Small 3 (2501) · 24B
GeLLM3O-Mistral
Ministral-3-14B-Base-2512 · 14B
Ministral-3-3B-Base-2512 · 3B
Ministral-3-8B-Base-2512 · 8B
Ministral 3 8B Reasoning · 8B
Mistral-7B-SFT-Beta · 7B
Mistral-7B-Instruct-v0.3 · 7B
Mistral NeMo 12B (base) · 12B
Mistral NeMo 12B Instruct (2407) · 12B
Mixtral-8x7B-Instruct-v0.1 · 8x7B
SFR-Embedding-Mistral · 7B

Observations (61)

Papers

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering (2026), On the Mutual Influence of Gender and Occupation in LLM Representations (2025), Concept Heterogeneity-aware Representation Steering (2026), Latent Structure of Affective Representations in Large Language Models (2026), A Minimalist Structural Probe for Phase-Count Abstraction in Transformer Language Models (2026), Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning (2026), Linear Representations of Political Perspective Emerge in Large Language Models (2025), Investigating Representation Universality: Case Study on Genealogical Representations (2024), Head Pursuit: Probing Attention Specialization in Multimodal Transformers (2025), Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models (2026), The Geometry of Categorical and Hierarchical Concepts in Large Language Models (2024), When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs (2026), Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness (2026), Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations (2026), Emergent Manifold Separability during Reasoning in Large Language Models (2026), Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations (2026), The Lattice Representation Hypothesis of Large Language Models (2026), The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models (2026), Language Models Represent and Transform Concepts with Shared Geometry (2026), Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs (2025), Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization (2026), The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces (2024), Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability (2026), How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? (2026), Emergent Causal-Geometric Dynamics Across Depth in Large Language Models (2026), Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries (2026), A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension (2024), Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing (2024), Geometric Signatures of Compositionality Across a Language Model's Lifetime (2024), The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning (2024), Rethinking Intrinsic Dimension Estimation in Neural Representations (2026), Geometric Asymmetry in MoE Specialization: Functional Decorrelation and Representational Overlap (2026), Language Models Encode Numbers Using Digit Representations in Base 10 (2025), Number Representations in LLMs: A Computational Parallel to Human Perception (2025), Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling (2025), Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models (2026), LLMs Know More About Numbers than They Can Say (2026), Relational Linearity is a Predictor of Hallucinations (2026), The Origins of Representation Manifolds in Large Language Models (2025), Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens (2026), How Language Models Process Negation (2026), Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks (2025), Learning Uncertainty from Sequential Internal Dispersion in Large Language Models (2026), A polar coordinate system represents syntax in large language models (2024), Model Editing as a Robust and Denoised Variant of DPO: A Case Study on Toxicity (2024), Emotions Where Art Thou: Characterizing the Emotional Latent Space of LLMs (2025), Intervention Lens: from Representation Surgery to String Counterfactuals (2024), The Platonic Representation Hypothesis (2024), Understanding and Preserving Safety in Fine-Tuned LLMs (2026), Language Models Are Implicitly Continuous (2025), Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention (2026), The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations (2026), Revisiting Hallucination Detection with Effective Rank-based Uncertainty (2025), Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings (2026), Lines of Thought in Large Language Models (2024), Large Language Models Encode Semantics and Alignment in Linearly Separable Representations (2025), Symmetry in Language Statistics Shapes the Geometry of Model Representations (2026), Not All Language Model Features Are One-Dimensionally Linear (2024), Do Sparse Autoencoders Capture Concept Manifolds? (2026), What are you sinking? A geometric approach on attention sink (2025), The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets (2024), On the Universal Truthfulness Hyperplane Inside LLMs (2024), How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs (2026), Tracing Moral Foundations in Large Language Models (2026), Revisiting the Othello World Model Hypothesis (2025), SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing (2026), Language Models Represent Beliefs of Self and Others (2024)