Correct-to-incorrect sycophancy localizes to sparse middle-layer heads
measured in 1 paperGenadi et al. train logistic-regression probes on residual, MLP, and per-head attention activations of Gemma-3-4B and Llama-3.2-3B to localize correct-to-incorrect sycophancy [genadi-etal-2026-sycophancy-hides-linearly-attention-heads] Residual and MLP probes are broadly accurate (Gemma residual 99.6%), but attention-probe accuracy concentrates sharply in a sparse subset of middle-layer heads [genadi-etal-2026-sycophancy-hides-linearly-attention-heads] Steering the probe direction at those heads cuts sycophancy from 40.7% to 34.4% (Gemma-3) and 51.7% to 25.0% (Llama-3.2), while MLP/residual steering underperforms [genadi-etal-2026-sycophancy-hides-linearly-attention-heads] The sycophancy direction is only mildly anti-correlated with a truthful direction (cosine -0.22, 32% head overlap), indicating related but distinct linear mechanisms [genadi-etal-2026-sycophancy-hides-linearly-attention-heads]