MATH · IN · MODELS

Steering vectors are non-identifiable; orthogonal perturbations are equivalent

measured in 1 paper

Venkatesh & Kurapath prove via a Jacobian null-space argument that steering vectors extracted for a target behavior are not unique [venkatesh-kurapath-2026-on-the-non-identifiability-of-steering-vectors-in-large-language-models] On Llama-3.1-8B-Instruct and Qwen2.5-3B-Instruct across five traits, orthogonal perturbations within the activation-covariance null space produce behavioral effects indistinguishable from the original (mean Cohen's d 0.12-0.13, below the 0.2 detectability threshold) [venkatesh-kurapath-2026-on-the-non-identifiability-of-steering-vectors-in-large-language-models] A geometrically distinct mean-difference vector and a PCA-derived vector for the same trait are behaviorally equivalent [venkatesh-kurapath-2026-on-the-non-identifiability-of-steering-vectors-in-large-language-models] Many different directions in representation space thus implement the same steering effect [venkatesh-kurapath-2026-on-the-non-identifiability-of-steering-vectors-in-large-language-models]

Context

steering, identifiability

Papers

On the Non-Identifiability of Steering Vectors in Large Language Models — Venkatesh, Sohan, Kurapath, Ashish Mahendran2026 · arXiv:2602.06801