MATH · IN · MODELS

Geographic representations are causally used for prediction, not just decodable

measured in 1 paper

Chen et al. extend the linear-geography finding to DeBERTa-v2-xxlarge and GPT-Neo-1.3B, where non-linear probes substantially outperform linear ones (unlike Gurnee & Tegmark) and a Haversine-based GeoDist loss yields probes whose predicted maps better resemble true geography [chen-etal-2023] RSA on country-level activation distances correlates modestly with real distance (Kendall tau 0.20-0.21), flagged by the authors as suggestive not conclusive [chen-etal-2023] Probe-gradient perturbation establishes causality: sharpening a city's alignment with its true coordinates improves country classification while gradient ascent degrades it, and cross-city perturbation shifts logits toward the target country [chen-etal-2023] Applied to GPT-Neo's own next-token prediction, gradient ascent causes a large drop in country-token accuracy and descent a small significant improvement, so the language-modeling objective itself depends on the spatial representation [chen-etal-2023] Scope: the models are smaller and older (both under 2B), which may limit how strongly the causal finding generalizes to frontier scale [chen-etal-2023]

Context

space, geography, world models, causal validation, representational similarity analysis, non-linear probing, geography representation debate

Papers

More than Correlation: Do Large Language Models Learn Causal Representations of Space? — Chen, Yida, Gan, Yixian, Li, Sijia, Yao, Li, Zhao, Xiaohan2023 · arXiv:2312.16257