A difficulty direction is language-agnostic shallow, language-specific deep
measured in 1 paperCivelli et al. train per-layer linear probes on four LLMs to predict a continuous problem-difficulty score across 21 languages [civelli-etal-2026-shared-geometry-difficulty-multilingual] Deep-layer probes reach high same-language accuracy (Llama-3.1-8B rho=0.822 at layer ~30) but generalize poorly across languages [civelli-etal-2026-shared-geometry-difficulty-multilingual] Shallow-layer probes carry a language-agnostic difficulty signal (cross-lingual rho=0.783 at layer ~16) [civelli-etal-2026-shared-geometry-difficulty-multilingual] Fixing the deep same-language-optimal layer costs 0.177 rho cross-lingually while the shallow transfer-optimal layer costs only 0.014 rho in-language; replicated in Qwen3-8B [civelli-etal-2026-shared-geometry-difficulty-multilingual] No causal steering is performed, so the claim rests on decodability alone [civelli-etal-2026-shared-geometry-difficulty-multilingual]