MATH · IN · MODELS

Human evaluation confirms per-layer style vectors steer emotion

measured in 1 paper

Diallo et al. extract per-layer style/emotion vectors as contrastive mean-activation differences and inject them across all layers of Alpaca and Llama-3-8B-Lexi-Uncensored [diallo-etal-2026-the-effectiveness-of-style-vectors-for-steering-llms-a-human-evaluation] The first large-scale human evaluation of activation steering (7,000+ ratings, 190 participants) shows moderate steering (lambda ~0.15) reliably shifts perceived emotion [diallo-etal-2026-the-effectiveness-of-style-vectors-for-steering-llms-a-human-evaluation] Effects are large for disgust (partial eta-squared 0.616) and fear (0.540) but minimal for surprise (0.042) [diallo-etal-2026-the-effectiveness-of-style-vectors-for-steering-llms-a-human-evaluation] Human ratings agree strongly with an automated classifier (mean r=0.776), and Llama-3 steers more consistently than Alpaca (p<0.001) [diallo-etal-2026-the-effectiveness-of-style-vectors-for-steering-llms-a-human-evaluation]

Context

steering, emotion

Papers

The Effectiveness of Style Vectors for Steering LLMs: A Human Evaluation — Diallo, Diaoulé, Dworatzyk, Katharina, Jentzsch, Sophie, Schütt, Peer, Theis, Sabine, Hecking, Tobias2026 · arXiv:2601.21505