MATH · IN · MODELS

Negation is a single linear direction decodable by layer 4

measured in 1 paper

Zhou et al. extract a single "not" direction from Llama-3.1-8B residual-stream states via PCA-for-reduction followed by linear discriminant analysis [zhou-etal-2026-how-language-models-process-negation] Positive and negative hidden states are approximately linearly separable by this one direction [zhou-etal-2026-how-language-models-process-negation] 10-fold cross-validated per-layer decoding accuracy reaches near-perfect by layer 4 [zhou-etal-2026-how-language-models-process-negation] The same linear-direction analysis is confirmed on Mistral-7B-v0.1 as a secondary model [zhou-etal-2026-how-language-models-process-negation]

Context

negation, semantic-directions

Papers

How Language Models Process Negation — Zhou, Zhejian, Zhou, Tianyi, Jia, Robin, May, Jonathan2026 · arXiv:2605.03052