MATH · IN · MODELS

Per-task truth probe directions are near-orthogonal with disjoint L1 support

measured in 1 paper

Azizian et al. train L2-logistic truth probes separately on seven QA/fact-verification tasks across Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and Phi-4-Mini-Instruct [azizian-etal-2026-truth-geometries-orthogonal-across-tasks] Pairwise cosine similarity between probe directions is consistently low (<0.5) and correlates with cross-task transfer AUROC (r=0.59) [azizian-etal-2026-truth-geometries-orthogonal-across-tasks] L1 probes share under 15% of their nonzero support for most task pairs, and the few high-overlap pairs are exactly those that generalize well [azizian-etal-2026-truth-geometries-orthogonal-across-tasks] Three attempts to recover a shared cross-task truth direction (joint training, subspace-constrained fit, mixture-of-probes) all fail to beat per-task probes, and no causal intervention is performed [azizian-etal-2026-truth-geometries-orthogonal-across-tasks]

Context

task-specific linear truth probes (7 QA/fact-verification tasks), probe-direction cosine similarity as a predictor of cross-task transfer (r=0.59), L1 sparse-probe support overlap vs. cross-task generalization, failed recovery of a shared direction (joint training, subspace-constrained fit, mixture-of-experts), t-SNE visualization of per-task hidden-state clustering

Papers

The Geometries of Truth Are Orthogonal Across Tasks — Azizian, Waiss, Kirchhof, Michael, Ndiaye, Eugene, Bethune, Louis, Klein, Michal, Ablin, Pierre, Cuturi, Marco2026 · arXiv:2506.08572