Negation is a single linear direction decodable by layer 4
measured in 1 paperZhou et al. extract a single "not" direction from Llama-3.1-8B residual-stream states via PCA-for-reduction followed by linear discriminant analysis [zhou-etal-2026-how-language-models-process-negation] Positive and negative hidden states are approximately linearly separable by this one direction [zhou-etal-2026-how-language-models-process-negation] 10-fold cross-validated per-layer decoding accuracy reaches near-perfect by layer 4 [zhou-etal-2026-how-language-models-process-negation] The same linear-direction analysis is confirmed on Mistral-7B-v0.1 as a secondary model [zhou-etal-2026-how-language-models-process-negation]
Structure
Context
negation, semantic-directions
Confirmed in models
Papers
How Language Models Process Negation — Zhou, Zhejian, Zhou, Tianyi, Jia, Robin, May, Jonathan