OLMo
Allen Institute for AI
Structures found in this family (8)
By model (5)
OLMo-7B · 7B
OLMo-1B · 1B
OLMo 3 7B · 7B
OLMo-7B-0724-hf · 7B
OLMo-7B-SFT · 7B
Observations (13)
Papers
On the Mutual Influence of Gender and Occupation in LLM Representations (2025), Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization (2026), Polar probe linearly decodes semantic structures from LLMs (2026), Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics (2026), Characterizing Linear Alignment Across Language Models (2026), What Really Controls Temporal Reasoning in LLMs: Tokenisation or Representation of Time? (2026), Output Vector Editing for Memorization Mitigation in Large Language Models (2026), Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers (2026), Tracing the Representation Geometry of Language Models from Pretraining to Post-training (2025), Refusal in Language Models Is Mediated by a Single Direction (2024), Representation Engineering: A Top-Down Approach to AI Transparency (2023), Programming Refusal with Conditional Activation Steering (2024), Refusal Direction is Universal Across Safety-Aligned Languages (2025), The Platonic Representation Hypothesis (2024), Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology (2026), Scale Determines Whether Language Models Organize Representation Geometry for Prediction (2026)