MATH · IN · MODELS

Negative and positive valence localize to distinct depths; a negative-locus direction steers valence

measured in 1 paper

Venkatesh performs topic-controlled activation patching across all layers on Llama-3.2-1B-Instruct, Qwen2.5-1.5B-Instruct, and Qwen2.5-3B-Instruct with a shared corrupted baseline [venkatesh-2026-negative-before-positive] Negative-valence processing localizes causally to 14-27% of model depth and positive-valence to 53-66%, a consistent depth asymmetry (p<2e-10), with a flip-test ruling out simple topic detection [venkatesh-2026-negative-before-positive] A diff-in-means valence direction extracted at the negative-locus layer, added to neutral prompts, produces a monotonic dose-dependent valence shift (Spearman rho>0.89) [venkatesh-2026-negative-before-positive] This is distinct from the 2D circumplex account, being a depth-asymmetric causal localization of two separate one-dimensional valence directions [venkatesh-2026-negative-before-positive]

Context

valence, depth asymmetry, activation patching, diff-in-means, causal steering, topic-controlled baseline

Papers

Negative Before Positive: Asymmetric Valence Processing in Large Language Models — Venkatesh, Sohan2026 · arXiv:2605.05653