MATH · IN · MODELS

A Gemma-2-2B probe direction detects and bidirectionally steers hallucination

measured in 1 paper

O'Neill et al. train a linear probe on layer-10 residual activations of Gemma-2-2B to separate hallucinated from faithful summary continuations [oneill-etal-2026-single-direction-of-truth-hallucination] As a detector it reaches F1=0.97-0.99 on XSUM/CNN-DM, beating Lookback Lens by 5-8 points, and F1=0.75 on the cue-free CONTRATALES benchmark [oneill-etal-2026-single-direction-of-truth-hallucination] The same normalized probe direction is patched additively into layer 10 with alpha swept over [-60,60] [oneill-etal-2026-single-direction-of-truth-hallucination] Positive scaling raises hallucination rate to 0.86 while cutting repetition below 0.05; negative scaling raises repetition to 0.84 and lowers hallucination to 0.35 [oneill-etal-2026-single-direction-of-truth-hallucination] This is a same-model self-steering result combining detection accuracy with a geometry-tied causal intervention [oneill-etal-2026-single-direction-of-truth-hallucination]

Context

hallucination-detection linear probe (Gemma-2-2B, layer 10), cross-benchmark F1 (news summarization vs. logical-contradiction detection), additive steering with the probe's own normalized direction (alpha in [-60,60]), bidirectional hallucination-rate/repetition-rate trade-off

Papers

A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations — O'Neill, Charles, Chalnev, Slava, Zhao, Chi Chi, Kirkby, Max, Jayasekara, Mudith2026 · arXiv:2507.23221