MATH · IN · MODELS

A persona-conditioned PCA subspace separates deception and sycophancy

measured in 1 paper

Mahadik & Skapars build one activation vector per persona (deceptive/honest, sycophantic/non-sycophantic) by mean-pooling layer-14 hidden states, following the Assistant Axis method [mahadik-skapars-2026-persona-coordinates-linear-probes] PCA on the centered persona vectors gives a PC1 that cleanly separates harmful from harmless personas for both behaviors, with the default assistant near the harmless cluster [mahadik-skapars-2026-persona-coordinates-linear-probes] PC1 and a diff-in-means persona direction transfer zero-shot as classifiers to 5 unseen deception and 5 unseen sycophancy datasets [mahadik-skapars-2026-persona-coordinates-linear-probes] Probes on top-3 persona-PC features improve cross-dataset AUROC transfer over raw-activation probes and over random-subspace and dataset-PCA controls, on Llama-3.2-3B and replicated on Llama-3-8B [mahadik-skapars-2026-persona-coordinates-linear-probes]

Context

persona-conditioned activation vectors (mean-pooled over personas/questions) as PCA input, following the Assistant Axis method, PC1 as an unsupervised harmful/harmless persona separator, zero-shot cross-dataset transfer of a persona-derived contrastive direction and PC1, persona-PC-projected probe features vs. raw-activation, random-subspace, and dataset-PCA control baselines, within-behavior dataset clustering differing between sycophancy (bimodal) and deception (no clear clusters)

Papers

Do Linear Probes Generalize Better in Persona Coordinates? — Mahadik, Prasad, Skapars, Adrians2026 · arXiv:2605.09391