A persona-conditioned PCA subspace separates deception and sycophancy
measured in 1 paperMahadik & Skapars build one activation vector per persona (deceptive/honest, sycophantic/non-sycophantic) by mean-pooling layer-14 hidden states, following the Assistant Axis method [mahadik-skapars-2026-persona-coordinates-linear-probes] PCA on the centered persona vectors gives a PC1 that cleanly separates harmful from harmless personas for both behaviors, with the default assistant near the harmless cluster [mahadik-skapars-2026-persona-coordinates-linear-probes] PC1 and a diff-in-means persona direction transfer zero-shot as classifiers to 5 unseen deception and 5 unseen sycophancy datasets [mahadik-skapars-2026-persona-coordinates-linear-probes] Probes on top-3 persona-PC features improve cross-dataset AUROC transfer over raw-activation probes and over random-subspace and dataset-PCA controls, on Llama-3.2-3B and replicated on Llama-3-8B [mahadik-skapars-2026-persona-coordinates-linear-probes]