MATH · IN · MODELS

Temporal knowledge drift is a linear direction orthogonal to correctness and uncertainty

measured in 1 paper

Elbadry et al. train L1-regularized probes on six instruction-tuned LLMs to detect temporal drift (whether a fact changed since training cutoff), reaching AUROC 0.83-0.95 versus 0.49-0.57 for output and correctness/uncertainty baselines [elbadry-etal-2026-geometry-of-forgetting-temporal-knowledge-drift-as-an-independent-axis-in-llm-representations] Five measures establish the drift direction is geometrically orthogonal to correctness and uncertainty probe directions (weight cosine <=0.136; INLP removal of 10 directions changes AUROC <=0.013) [elbadry-etal-2026-geometry-of-forgetting-temporal-knowledge-drift-as-an-independent-axis-in-llm-representations] Untrained difference-of-means directions overlap with correctness/uncertainty, yet the trained regularized drift probe is near-orthogonal, so the orthogonality is a genuine concept property not a training artifact [elbadry-etal-2026-geometry-of-forgetting-temporal-knowledge-drift-as-an-independent-axis-in-llm-representations] A diff-in-means steering direction is silent under ablation but under amplification produces structured logit redistribution favoring the current over the stale fact holder (selectivity -2.46 to -10.55 logits) [elbadry-etal-2026-geometry-of-forgetting-temporal-knowledge-drift-as-an-independent-axis-in-llm-representations] A cross-cutoff entity-matched control (0.975-0.998 across seven model pairs) confirms the probe reads model-internal knowledge state, not an input property [elbadry-etal-2026-geometry-of-forgetting-temporal-knowledge-drift-as-an-independent-axis-in-llm-representations]

Context

temporal-drift linear probe (six instruction-tuned LLMs, entity-disjoint controlled AUROC 0.83-0.95), five-measure geometric-independence protocol (weight cosine, score correlation, null-space projection, INLP, DoM dissociation), orthogonality to correctness and uncertainty probe directions, diff-in-means steering direction (latent-but-activatable -- silent under ablation, structured under amplification), cross-cutoff entity-matched control isolating model-internal state from input properties

Papers

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations — Elbadry, Rania, Heakl, Ahmed, Zhang, Fan, Bouch, Dani, Wang, Yuxia, Nakov, Preslav, Xie, Zhuohan2026 · arXiv:2605.09195