Linear Artificial Tomography reads and controls concepts across LLMs
measured in 1 paperZou et al. introduce Linear Artificial Tomography (LAT): from contrastive stimulus pairs, take the first principal component of paired activation differences as a concept reading vector [zou-etal-2023] Across LLaMA-2-Chat (7B/13B/70B), Vicuna-13B, Vicuna-33B-Uncensored, and DeBERTa, a single direction classifies and causally controls honesty, ethics, morality, emotion, and bias [zou-etal-2023] LAT recovers truthfulness on DeBERTa more accurately than contrast-consistent search [zou-etal-2023] A harmfulness direction in Vicuna-13B stays a >90% classifier under jailbreaks, and boosting its salience raises harmless-response rates under attack [zou-etal-2023]
Structure
Context
honesty, truthfulness, utility, morality, power aversion, emotion, bias and fairness, representation engineering
Confirmed in models
Method
Papers
Representation Engineering: A Top-Down Approach to AI Transparency — Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A., Goel, S., Li, N., Byun, M., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, Z., Hendrycks, D.