Persona/style control localizes to a sparse set of attention heads
measured in 1 paperIzawa et al. localize persona/style control in Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct to 3 attention heads per model via layer-wise persona-vector heatmaps and a head-wise Head Contribution Score [izawa-etal-2026-steering-at-the-source-style-modulation-heads-for-robust-persona-control] Steering only these heads achieves the best trait-expression-vs-coherency Pareto frontier (best score in 11/12 Qwen conditions, 9/12 Llama) [izawa-etal-2026-steering-at-the-source-style-modulation-heads-for-robust-persona-control] Zero-ablating them causes a sharp targeted drop in trait expression while leaving MMLU and output coherency intact [izawa-etal-2026-steering-at-the-source-style-modulation-heads-for-robust-persona-control]
Structure
Context
persona-control, attention-head-localization
Confirmed in models
Papers
Steering at the Source: Style Modulation Heads for Robust Persona Control — Izawa, Yoshihiro, Minegishi, Gouki, Eguchi, Koshi, Hosokawa, Sosuke, Taura, Kenjiro