MATH · IN · MODELS

Big Five trait directions are near-orthogonal but steer only forced-choice

measured in 1 paper

Frising & Balcells fit OLS trait directions from 406 role-play character descriptions in Llama-3.3-70B-Instruct, separately per layer and token position [frising-balcells-2025-linear-personality-probing-steering-big-five] The five OCEAN regression directions show low cross-talk / near-orthogonality, unlike top-variance SVD directions that collapse toward one shared personality axis [frising-balcells-2025-linear-personality-probing-steering-big-five] Same-trait regression directions across token positions are well-aligned, the opposite of SVD's per-position behavior [frising-balcells-2025-linear-personality-probing-steering-big-five] Steering the mean-input-prompt direction monotonically shifts forced-choice Extraversion for |alpha|<=0.4, then degrades into gibberish [frising-balcells-2025-linear-personality-probing-steering-big-five] The steering effect vanishes once character context is already in the prompt, showing the direction's causal reach is context-dependent [frising-balcells-2025-linear-personality-probing-steering-big-five]

Context

approximate orthogonality between five independently-regressed trait directions, contrasting unsupervised (SVD/max-variance) baseline collapsing traits onto one shared axis, forced-choice vs. open-ended vs. context-present steering conditions, steering coefficient range (|alpha| <= 0.4) before output degradation

Papers

Linear Personality Probing and Steering in LLMs: A Big Five Study — Frising, Michel, Balcells, Daniel2025 · arXiv:2512.17639