MATH · IN · MODELS

Correct-to-incorrect sycophancy localizes to sparse middle-layer heads

measured in 1 paper

Genadi et al. train logistic-regression probes on residual, MLP, and per-head attention activations of Gemma-3-4B and Llama-3.2-3B to localize correct-to-incorrect sycophancy [genadi-etal-2026-sycophancy-hides-linearly-attention-heads] Residual and MLP probes are broadly accurate (Gemma residual 99.6%), but attention-probe accuracy concentrates sharply in a sparse subset of middle-layer heads [genadi-etal-2026-sycophancy-hides-linearly-attention-heads] Steering the probe direction at those heads cuts sycophancy from 40.7% to 34.4% (Gemma-3) and 51.7% to 25.0% (Llama-3.2), while MLP/residual steering underperforms [genadi-etal-2026-sycophancy-hides-linearly-attention-heads] The sycophancy direction is only mildly anti-correlated with a truthful direction (cosine -0.22, 32% head overlap), indicating related but distinct linear mechanisms [genadi-etal-2026-sycophancy-hides-linearly-attention-heads]

Context

correct-to-incorrect sycophancy, attention-head localization, probe-derived steering direction, truthful vs sycophancy direction dissociation

Papers

Sycophancy Hides Linearly in the Attention Heads — Genadi, Rifo, Nwadike, Munachiso, Mukhituly, Nurdaulet, Alquabeh, Hilal, Hiraoka, Tatsuya, Inui, Kentaro2026 · arXiv:2601.16644