Geographic representations are causally used for prediction, not just decodable
measured in 1 paperChen et al. extend the linear-geography finding to DeBERTa-v2-xxlarge and GPT-Neo-1.3B, where non-linear probes substantially outperform linear ones (unlike Gurnee & Tegmark) and a Haversine-based GeoDist loss yields probes whose predicted maps better resemble true geography [chen-etal-2023] RSA on country-level activation distances correlates modestly with real distance (Kendall tau 0.20-0.21), flagged by the authors as suggestive not conclusive [chen-etal-2023] Probe-gradient perturbation establishes causality: sharpening a city's alignment with its true coordinates improves country classification while gradient ascent degrades it, and cross-city perturbation shifts logits toward the target country [chen-etal-2023] Applied to GPT-Neo's own next-token prediction, gradient ascent causes a large drop in country-token accuracy and descent a small significant improvement, so the language-modeling objective itself depends on the spatial representation [chen-etal-2023] Scope: the models are smaller and older (both under 2B), which may limit how strongly the causal finding generalizes to frontier scale [chen-etal-2023]