Sycophantic agreement, genuine agreement, and praise occupy distinct steerable directions
measured in 1 paperVennemeyer et al. decompose sycophancy into sycophantic agreement, genuine agreement, and sycophantic praise, extracting a diff-in-means direction per behavior at every layer over items where the model knows the ground truth [vennemeyer-etal-2025-sycophancy-not-one-thing] Sycophantic- and genuine-agreement directions start nearly identical (cosine ~0.99 early) then diverge sharply to ~0.07 by layer 25, while the praise direction stays near-orthogonal to both throughout [vennemeyer-etal-2025-sycophancy-not-one-thing] Ablating a behavior's own diff-in-means subspace drops its own detection AUROC to chance while removing the praise subspace has zero effect on agreement detection [vennemeyer-etal-2025-sycophancy-not-one-thing] Activation-addition steering is highly selective (praise 36.8x in LLaMA-3.1-8B, sycophantic-agreement 23.1x in Qwen3-30B), replicated across five models [vennemeyer-etal-2025-sycophancy-not-one-thing]