MATH · IN · MODELS

Steering directions split into a magnitude axis and interference plane

measured in 1 paper

Gao et al. relax the Linear Representation Hypothesis orthogonality assumption, decomposing a concept steering vector into a central axis and an orthogonal 2D normal plane spanned by the axis complement and the top PC of other concepts directions [gao-etal-2026-cylindrical-hypothesis] Despite cylindrical/phase/sector terminology there is no measured angular or periodic structure (phase = position in the 2D plane; sectors = a binary high/low-sensitivity split), so it is a linear-subspace interference decomposition, not a topology [gao-etal-2026-cylindrical-hypothesis] Steering-effect magnitude follows a predictable sin^m*cos^n form, but which interference sector a concept pair falls into is NOT predictable from the vectors (Pearson -0.034), a genuine null [gao-etal-2026-cylindrical-hypothesis] A penalty experiment attenuating the plane component trades earlier target onset against earlier corrupted output, validated by an LLM-judge at 94% human agreement [gao-etal-2026-cylindrical-hypothesis]

Context

concept interference, steering predictability, LRH orthogonality relaxation, sensitivity sectors, causal steering trade-off, LLM-judge validation

Papers

The Cylindrical Representation Hypothesis for Language Model Steering — Gao, Lang, Zhang, Jinghui, Liu, Wei, Ji, Fengxian, Wang, Chenxi, Song, Zirui, Ghosh, Akash, Mohamed, Youssef, Nakov, Preslav, Chen, Xiuying2026 · arXiv:2605.01844