Persona-vector taxonomy sorts traits as natural, steerable, or intractable
measured in 1 paperZeng, Emami & Choi extract 53 diff-in-means contrastive persona vectors across four trait domains in Qwen3-8B and gpt-oss-20b [zeng-etal-2026-what-models-express-suppress-and-resist-auditing-open-weight-llms-with-persona-vectors] Sweeping steering strength alpha in {0,0.5,...,2.5} classifies each trait as natural, steerable, or intractable via quantitative thresholds (baseline>=70 and gain<=10 for "natural") [zeng-etal-2026-what-models-express-suppress-and-resist-auditing-open-weight-llms-with-persona-vectors] Mean steering-gain is 18.39 (clinician domain, Qwen3-8B) and 11.56 (gpt-oss-20b), and all 9 agentic traits are natural in both models [zeng-etal-2026-what-models-express-suppress-and-resist-auditing-open-weight-llms-with-persona-vectors] Clinician defaults matched a board-certified psychiatrist's desirability judgment on 16 of 17 traits [zeng-etal-2026-what-models-express-suppress-and-resist-auditing-open-weight-llms-with-persona-vectors] Across 171 pairwise generic-trait steering combinations, destructive composition requires two steerable traits and never occurs when a natural trait is involved [zeng-etal-2026-what-models-express-suppress-and-resist-auditing-open-weight-llms-with-persona-vectors] For the intractable "evil" trait in gpt-oss-20b, a persona vector transferred from a fine-tuned variant recovers behavior the base model refuses to express [zeng-etal-2026-what-models-express-suppress-and-resist-auditing-open-weight-llms-with-persona-vectors]