MATH · IN · MODELS

Persona-vector projections predict and prevent trait drift

measured in 1 paper

Chen, Arditi, Sleight, Evans & Lindsey extract persona vectors (diff of mean response-token activations) for evil, sycophancy, and hallucination in Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct [chen-etal-2025-persona-vectors] Final-prompt-token projections correlate with subsequent trait expression (Pearson r=0.634-0.830; 94.7% judge agreement) [chen-etal-2025-persona-vectors] The same projection over a finetuning dataset's responses predicts how much that run shifts trait propensity (r=0.76-0.97), flagging problematic data before finetuning [chen-etal-2025-persona-vectors] Proactively steering toward the undesired persona during training reduces trait shifts while keeping coherence above 80 and better preserving MMLU than regular finetuning [chen-etal-2025-persona-vectors]

Context

persona-vector projection correlates with trait expression, dataset-level projection flags problematic training data, preventative training-time steering

Papers

Persona Vectors: Monitoring and Controlling Character Traits in Language Models — Chen, Runjin, Arditi, Andy, Sleight, Henry, Evans, Owain, Lindsey, Jack2025 · arXiv:2507.21509