A mid-layer head group decodes another agent's beliefs and steers theory-of-mind
measured in 1 paperZhu et al. prompt Mistral-7B-Instruct (and DeepSeek-LLM-7B-Chat) with third-person belief narratives and fit logistic-regression probes on attention-head activations [zhu-etal-2024-language-models-represent-beliefs-of-self-and-others] A group of middle-layer attention heads decodes another agent's belief status at over 80% accuracy, distinct from the model's own self-belief representation [zhu-etal-2024-language-models-represent-beliefs-of-self-and-others] Steering these representations produces large changes in downstream Theory-of-Mind performance, generalizing across social-reasoning tasks [zhu-etal-2024-language-models-represent-beliefs-of-self-and-others]
Structure
Context
dissociable self-vs-other belief representations recovered via separate linear probes on the same model's activations, causal steering of a belief representation validated against downstream behavioral task performance, not just probe accuracy
Confirmed in models
Papers
Language Models Represent Beliefs of Self and Others — Zhu, Wentao, Zhang, Zhining, Wang, Yizhou