A valence-arousal circumplex subspace causally steers emotion and refusal
measured in 1 paperSun et al. recover two near-orthogonal valence/arousal axes inside Llama-3.1-8B-Instruct (replicated in Qwen3-8B/14B) by PCA-projecting contrastive emotion-steering vectors and ridge-regressing against human ratings [sun-etal-2026-valence-arousal] Projecting the emotion vectors onto this plane traces a circle: a circularity statistic reaches 2.76-4.08 with fitted radii ~0.37-0.39, analogous to Russell's circumplex [sun-etal-2026-valence-arousal] The valence axis recovers self-reported valence at r=0.97 and the NRC-VAD lexicon at r=0.71, with cross-model valence agreement r=0.95 [sun-etal-2026-valence-arousal] Adding valence/arousal directions at specific circle angles produces dose-dependent, angle-specific shifts in generated-text affect (e.g. 0deg: delta-valence +0.75; 180deg: -0.73) [sun-etal-2026-valence-arousal] The same arousal axis causally controls refusal (20%->86% on OKTest) and sycophancy, with random-direction controls within 2-3 points of baseline [sun-etal-2026-valence-arousal] Logit-clamping and top-neuron ablation along the direction crash refusal while preserving MATH-500/IFEval, and an independent refusal direction is near-orthogonal (86.5deg) to the VA plane [sun-etal-2026-valence-arousal] Van der Ben et al. independently replicate the valence/arousal PCA structure (PC1-valence r=0.72-0.83, PC2-arousal r=0.21-0.45) in Apertus-8B and Gemma-4-E4B-it without computing circularity or steering [vanderben-etal-2026-emotion-vectors-open-source-llms] They add that cross-architecture layer-depth trajectories diverge sharply (a 3-phase plateau in Apertus vs a smooth gradient in Gemma via linear CKA) [vanderben-etal-2026-emotion-vectors-open-source-llms]