MATH · IN · MODELS

Numbers lie on a logarithmically-compressed 1D number line

measured in 1 paper

Alquboj et al. project last-token numeric representations across six base models onto the top 1-2 PCA components, revealing a monotonic open number line whose spacing shrinks with magnitude (Scaling Rate Index beta<1), consistent with a logarithmic scale [alquboj-etal-2025] A target-aware PLS probe does not surface this compression while PCA's target-blind projection does, and the same compressed structure appears for birth years (not populations) in Llama-3 instruct models [alquboj-etal-2025] Tiblias et al. extend log-compression to temporal reasoning: durations and recurrence frequencies are best fit by a log-linear/log-semicircular manifold across Qwen2.5-3B, Llama-3.2-3B, and Gemma-2-2B, persisting to 70B scale [tiblias-etal-2025] Cacioli confirms log-compression across numerical, temporal, and spatial domains and three model families, with Weber R^2 up to .83-.88 versus linear .14-.32 [cacioli-2026-webers-law] A magnitude direction (a ridge probe, also the dominant PCA axis at r>.80 with log-magnitude) causally and dose-dependently shifts magnitude-comparison behavior at early layers (4.1x over random controls) [cacioli-2026-webers-law] The causal effect is layer-dissociated from geometry strength: at later layers where the structure is most pronounced the same intervention is inert (1.2x), and Mistral shows the geometry but not a human-range Weber fraction [cacioli-2026-webers-law] Yuchi et al. confirm log-magnitude compression via a supervised ridge probe across seven models, recovering numbers at ~2.3% relative error (rho=0.71 on real arXiv text) [yuchi-etal-2026-llms-know-more-about-numbers] They add a geometry-behavior dissociation: the hidden state linearly decodes a numeral pair's ranking at over 90% while verbalized comparison reaches only 50-70%, and an auxiliary-probe fine-tune closes this articulation gap by up to 84% [yuchi-etal-2026-llms-know-more-about-numbers]

Context

numbers, logarithmic mental number line, sublinear compression, order preservation, birth years, human cognition parallel, temporal reasoning, task-dependent geometry, causally validated magnitude direction, geometry-behavior dissociation, layer-specific causal locus

Papers

Number Representations in LLMs: A Computational Parallel to Human Perception — AlquBoj, H. V., AlQuabeh, H., Bojkovic, V., Hiraoka, T., El-Shangiti, A. O., Nwadike, M., Inui, K.2025 · arXiv:2502.16147
Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling — Tiblias, F., Bigoulaeva, I., Niu, J., Balloccu, S., Gurevych, I.2025 · arXiv:2510.01025
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models — Cacioli, Jon-Paul2026 · arXiv:2603.20642
LLMs Know More About Numbers than They Can Say — Yuchi, Fengting, Du, Li, Eisner, Jason2026 · arXiv:2602.07812