Residual-stream probes detect knowledge conflict and reliance
measured in 1 paper- Logistic-regression probes on the mid-layer residual stream detect parametric-vs-contextual knowledge conflict with ~90% accuracy (peaking around layer 14 of Llama-3-8B on NQSwap). [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts] - The same probes predict which knowledge source the model will rely on before generation (peaking around layers 16-17). [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts] - When the model relies on contextual knowledge the residual stream is distinctly more skewed than for parametric knowledge (most pronounced in layers ~20-30). [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts] - Tested on base Llama-3-8B and Llama-2-7B over NQSwap, Macnoise and ConflictQA; observational. [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts]