MATH · IN · MODELS

Concept information is angular, yet norm-preservation is not optimal steering

measured in 1 paper

Aparin & Gaintseva decompose each hidden state into a radial norm and an angular concept score against a unit steering direction, across seven models [aparin-gaintseva-2026-geometric-account-activation-steering-angle-norm-decomposition] Linear probes on normalized hidden states match raw probes while norm-only probes stay near chance, so concept information is essentially angular, not radial [aparin-gaintseva-2026-geometric-account-activation-steering-angle-norm-decomposition] They systematically compare six steering variants that vary norm-preservation and angular-target enforcement [aparin-gaintseva-2026-geometric-account-activation-steering-angle-norm-decomposition] Despite concepts being angular, strict norm preservation is not most stable: moving the radial scale from beta=1.0 to 1.2 improves perplexity ~1.8x at a task-metric cost within ~2.5 points [aparin-gaintseva-2026-geometric-account-activation-steering-angle-norm-decomposition] This dissociates where a concept lives (angle) from what a stable intervention should manipulate (angle plus a non-unit radial scale) [aparin-gaintseva-2026-geometric-account-activation-steering-angle-norm-decomposition]

Context

angle-norm decomposition of a hidden state relative to a steering direction, linear probing on raw vs. normalized vs. norm-only representations to isolate where concept information lives, a systematic multi-axis comparison of steering variants tied directly to the angle-norm decomposition

Papers

A Geometric Account of Activation Steering through Angle-Norm Decomposition — Aparin, Georgii, Gaintseva, Tatiana2026 · arXiv:2606.06735