Temporal-probe decodability predicts reasoning in high-resource languages
measured in 1 paperBhatia et al. introduce MultiTempBench and evaluate 20 LLMs, combining a tokenization-fragmentation metric with linear-regression probes decoding Year/Month/Day (probe R^2 as temporal linearity) [bhatia-etal-2026-what-really-controls-temporal-reasoning-in-llms] A crossed mixed-effects regression over 285,000 predictions finds temporal linearity is the strongest predictor of reasoning accuracy in high-resource languages (English r=0.77) [bhatia-etal-2026-what-really-controls-temporal-reasoning-in-llms] Tokenization fragmentation dominates in low-resource languages (Hausa r=-0.97); the finding is purely correlational [bhatia-etal-2026-what-really-controls-temporal-reasoning-in-llms]
Structure
Context
temporal-reasoning, probe-geometry
Confirmed in models
Method
Papers
What Really Controls Temporal Reasoning in LLMs: Tokenisation or Representation of Time? — Bhatia, Gagan, Isa, Ahmad Muhammad, Peyrard, Maxime, Zhao, Wei