MATH · IN · MODELS

Whisper hallucination is linearly decodable and causally steerable

measured in 1 paper

Aparin et al. fit per-layer logistic classifiers to detect hallucination from frozen Whisper small/large-v3 encoder activations and from Batch-TopK SAE latents [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] Both representation spaces are linearly separable for hallucination status (layer-averaged AUC 0.74-0.80), improving toward deeper layers [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] Hallucination information concentrates in a small subset of SAE features, stabilizing at 50-100 features [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] A diff-in-means direction added to the final encoder layer, and a sparse sign-pattern over top SAE features, both causally steer hallucination [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] SAE steering cuts hallucination rate from 72.63% to 14.11% (small) and 86.88% to 27.33% (large-v3), at a modest clean-speech WER cost [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders]

Context

hallucination status as a linearly decodable property with a layer-depth trend, in both raw and SAE representation spaces, concentration of a behavioral property in a small subset of a much larger sparse feature bank, additive causal steering of both a diff-in-means direction and a sparse SAE sign-pattern direction, quantified against a shared behavioral metric

Papers

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders — Aparin, Georgii, Popov, Vadim, Sadekova, Tasnima, Yermekova, Assel2026 · arXiv:2606.07473