Cue-induced bias directions are installed by alignment tuning
measured in 1 paperGupta et al. extract a per-bias diff-of-means direction for seven cue-triggered bias types across base and instruct checkpoints of five model families [gupta-etal-2026-alignment-tuning-sycophancy-directions] Four of five base models flip on under 5% as many pairs as their instruct counterparts, with base activations carrying essentially no cue-specific signal, so susceptibility is installed by alignment tuning [gupta-etal-2026-alignment-tuning-sycophancy-directions] Within instruct models each bias direction transfers to held-out data at mean AUROC 0.69-0.82, localizing to late-middle layers [gupta-etal-2026-alignment-tuning-sycophancy-directions] Biases stay representationally distinct, with even behaviorally similar biases occupying different, sometimes anti-aligned directions [gupta-etal-2026-alignment-tuning-sycophancy-directions] Subtracting the direction while preserving over 90% of correct answers recovers 7-20% of bias-induced errors, versus under 5% for a random direction [gupta-etal-2026-alignment-tuning-sycophancy-directions]