Negative and positive valence localize to distinct depths; a negative-locus direction steers valence
measured in 1 paperVenkatesh performs topic-controlled activation patching across all layers on Llama-3.2-1B-Instruct, Qwen2.5-1.5B-Instruct, and Qwen2.5-3B-Instruct with a shared corrupted baseline [venkatesh-2026-negative-before-positive] Negative-valence processing localizes causally to 14-27% of model depth and positive-valence to 53-66%, a consistent depth asymmetry (p<2e-10), with a flip-test ruling out simple topic detection [venkatesh-2026-negative-before-positive] A diff-in-means valence direction extracted at the negative-locus layer, added to neutral prompts, produces a monotonic dose-dependent valence shift (Spearman rho>0.89) [venkatesh-2026-negative-before-positive] This is distinct from the 2D circumplex account, being a depth-asymmetric causal localization of two separate one-dimensional valence directions [venkatesh-2026-negative-before-positive]