MATH · IN · MODELS

Attention-head probes read and steer political ideology

measured in 1 paper

Kim, Evans & Schein fit a ridge-regression probe per attention head in three open chat LLMs to predict 552 U.S. lawmakers' DW-NOMINATE ideology scores [kim-evans-schein-2025-political-perspective] The best single head reaches Spearman rho 0.846-0.861, and an ensemble of the top 32 heads reaches 0.870-0.885, concentrated in middle layers [kim-evans-schein-2025-political-perspective] A nonlinear MLP probe matches the linear probe, supporting a linear-direction characterization [kim-evans-schein-2025-political-perspective] Probes fit on lawmaker ideology transfer zero-shot to predicting 400 news outlets' slant (rho 0.720-0.798) [kim-evans-schein-2025-political-perspective] Adding scaled top-head directions to activations shifts GPT-4o-rated political slant, correlating up to 0.607 with steering magnitude [kim-evans-schein-2025-political-perspective]

Context

per-attention-head ridge-regression probing (1,024 probes/model), DW-NOMINATE ideology score as a continuous linear target, zero-shot probe transfer to news-outlet slant, additive multi-head steering with GPT-4o + human-validated slant rating, middle layers most predictive

Papers

Linear Representations of Political Perspective Emerge in Large Language Models — Kim, Junsol, Evans, James, Schein, Aaron2025 · arXiv:2503.02080