MATH · IN · MODELS

Hyperplane-perturbation editing shifts personality with minimal capability loss

measured in 1 paper

Ju et al. convert each layer's personality linear probe into an editing mechanism, adding a closed-form correction along the probe's weight direction only when the probe misses the target trait [ju-etal-2025-personality-editing] The push is calibrated to just cross the decision boundary at a target confidence, unlike a fixed-magnitude steering vector applied uniformly [ju-etal-2025-personality-editing] Against in-context (IKE) and weight-editing (MEND) baselines on LLaMA-2-7B-Chat, LLaMA-3.1-8B-Instruct, and LLaMA-3.1-8B-Base across six trait conversions, it achieves substantially higher success (average 44.33%) [ju-etal-2025-personality-editing] The calibrated single-step edit preserves general capabilities better than the baselines [ju-etal-2025-personality-editing]

Context

closed-form boundary-crossing perturbation, personality-alignment-error (PAE) metric, trait-conversion difficulty tracking middle-layer geometry

Papers

Probing then Editing Response Personality of Large Language Models — Ju, Tianjie, Shao, Zhenyu, Wang, Bowen, Chen, Yujia, Zhang, Zhuosheng, Fei, Hao, Lee, Mong-Li, Hsu, Wynne, Duan, Sufeng, Liu, Gongshen2025 · arXiv:2504.10227