A Gemma-2-2B probe direction detects and bidirectionally steers hallucination
measured in 1 paperO'Neill et al. train a linear probe on layer-10 residual activations of Gemma-2-2B to separate hallucinated from faithful summary continuations [oneill-etal-2026-single-direction-of-truth-hallucination] As a detector it reaches F1=0.97-0.99 on XSUM/CNN-DM, beating Lookback Lens by 5-8 points, and F1=0.75 on the cue-free CONTRATALES benchmark [oneill-etal-2026-single-direction-of-truth-hallucination] The same normalized probe direction is patched additively into layer 10 with alpha swept over [-60,60] [oneill-etal-2026-single-direction-of-truth-hallucination] Positive scaling raises hallucination rate to 0.86 while cutting repetition below 0.05; negative scaling raises repetition to 0.84 and lowers hallucination to 0.35 [oneill-etal-2026-single-direction-of-truth-hallucination] This is a same-model self-steering result combining detection accuracy with a geometry-tied causal intervention [oneill-etal-2026-single-direction-of-truth-hallucination]