MATH · IN · MODELS

Sentiment is a single convergent linear direction, causally validated

measured in 1 paper

Tigges et al. extract a candidate sentiment direction in GPT2-small and Pythia (1.4B, 2.8B) via five techniques (mean difference, k-means, logistic regression, PCA, and DAS) [tigges-etal-2023-linear-sentiment] Pairwise cosine similarities between the five directions reach 72.6-99.1% (random baseline 0-2.4%), evidence they locate the same underlying direction [tigges-etal-2023-linear-sentiment] Directional ablation of the DAS direction on SST drops accuracy 100% to 62% (a 71% logit-difference reduction), and steering at coefficient -17 makes GPT2-small completions extremely negative [tigges-etal-2023-linear-sentiment] Single-scalar projection classifies token sentiment at 78-89%, and increasing the DAS subspace dimension does not improve OOD generalization, so one-dimensionality is treated as a supported but not final hypothesis [tigges-etal-2023-linear-sentiment]

Context

convergent validity across direction-extraction methods, directional activation patching, directional ablation, dimensionality-of-hypothesized-subspace caveat

Papers

Linear Representations of Sentiment in Large Language Models — Tigges, Curt, Hollinsworth, Oskar John, Geiger, Atticus, Nanda, Neel2023 · arXiv:2310.15154