Persona-vector projections predict and prevent trait drift
measured in 1 paperChen, Arditi, Sleight, Evans & Lindsey extract persona vectors (diff of mean response-token activations) for evil, sycophancy, and hallucination in Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct [chen-etal-2025-persona-vectors] Final-prompt-token projections correlate with subsequent trait expression (Pearson r=0.634-0.830; 94.7% judge agreement) [chen-etal-2025-persona-vectors] The same projection over a finetuning dataset's responses predicts how much that run shifts trait propensity (r=0.76-0.97), flagging problematic data before finetuning [chen-etal-2025-persona-vectors] Proactively steering toward the undesired persona during training reduces trait shifts while keeping coherence above 80 and better preserving MMLU than regular finetuning [chen-etal-2025-persona-vectors]