MATH · IN · MODELS

Sycophantic agreement, genuine agreement, and praise occupy distinct steerable directions

measured in 1 paper

Vennemeyer et al. decompose sycophancy into sycophantic agreement, genuine agreement, and sycophantic praise, extracting a diff-in-means direction per behavior at every layer over items where the model knows the ground truth [vennemeyer-etal-2025-sycophancy-not-one-thing] Sycophantic- and genuine-agreement directions start nearly identical (cosine ~0.99 early) then diverge sharply to ~0.07 by layer 25, while the praise direction stays near-orthogonal to both throughout [vennemeyer-etal-2025-sycophancy-not-one-thing] Ablating a behavior's own diff-in-means subspace drops its own detection AUROC to chance while removing the praise subspace has zero effect on agreement detection [vennemeyer-etal-2025-sycophancy-not-one-thing] Activation-addition steering is highly selective (praise 36.8x in LLaMA-3.1-8B, sycophantic-agreement 23.1x in Qwen3-30B), replicated across five models [vennemeyer-etal-2025-sycophancy-not-one-thing]

Context

diff-in-means direction per behavior, tracked by cosine similarity across layers, layer-dependent divergence of two initially near-identical directions (0.99 to 0.07 cosine), selectivity ratio (primary steering effect / largest cross-behavior effect) as a causal-independence metric, subspace-removal ablation isolating each behavior's own detectability, cross-model-family and cross-scale replication of both the geometric and causal pattern

Papers

Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs — Vennemeyer, Daniel, Duong, Phan Anh, Zhan, Tiffany, Jiang, Tianyu2025 · arXiv:2509.21305