MATH · IN · MODELS

Personality traits are linearly encoded with a betweenness arrangement

measured in 1 paper

Ju et al. train layer-wise linear probes on 11 open LLMs (including InternLM-7B and InternLM2-7B, base and instruct) to classify which of three personality instructions a response is conditioned on, scored by V-information [ju-etal-2025-personality-editing] V-information rises through layers 5-12 then plateaus, with instruction-tuned models 1-3 points higher than their base counterparts [ju-etal-2025-personality-editing] A t-SNE of Llama-3.1-8B-Instruct shows Neuroticism and Agreeableness distant with Extraversion situated between them, an informal betweenness arrangement [ju-etal-2025-personality-editing] This geometry explains why converting either trait into Extraversion is easier than converting directly between Neuroticism and Agreeableness [ju-etal-2025-personality-editing]

Context

relative arrangement (betweenness) of multiple trait directions, V-information as a layer-wise fit metric, instruction-tuning's effect on concept separability

Papers

Probing then Editing Response Personality of Large Language Models — Ju, Tianjie, Shao, Zhenyu, Wang, Bowen, Chen, Yujia, Zhang, Zhuosheng, Fei, Hao, Lee, Mong-Li, Hsu, Wynne, Duan, Sufeng, Liu, Gongshen2025 · arXiv:2504.10227