MATH · IN · MODELS

Residual-stream probes detect knowledge conflict and reliance

measured in 1 paper

- Logistic-regression probes on the mid-layer residual stream detect parametric-vs-contextual knowledge conflict with ~90% accuracy (peaking around layer 14 of Llama-3-8B on NQSwap). [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts] - The same probes predict which knowledge source the model will rely on before generation (peaking around layers 16-17). [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts] - When the model relies on contextual knowledge the residual stream is distinctly more skewed than for parametric knowledge (most pronounced in layers ~20-30). [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts] - Tested on base Llama-3-8B and Llama-2-7B over NQSwap, Macnoise and ConflictQA; observational. [zhao-etal-2024-analysing-the-residual-stream-of-language-models-under-knowledge-conflicts]

Context

knowledge-conflict, residual-stream-geometry

Papers

Analysing the Residual Stream of Language Models Under Knowledge Conflicts — Zhao, Yu, Du, Xiaotang, Hong, Giwon, Gema, Aryo Pradipta, Devoto, Alessio, Wang, Hongru, He, Xuanli, Wong, Kam-Fai, Minervini, Pasquale2024 · arXiv:2410.16090