Whisper hallucination is linearly decodable and causally steerable
measured in 1 paperAparin et al. fit per-layer logistic classifiers to detect hallucination from frozen Whisper small/large-v3 encoder activations and from Batch-TopK SAE latents [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] Both representation spaces are linearly separable for hallucination status (layer-averaged AUC 0.74-0.80), improving toward deeper layers [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] Hallucination information concentrates in a small subset of SAE features, stabilizing at 50-100 features [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] A diff-in-means direction added to the final encoder layer, and a sparse sign-pattern over top SAE features, both causally steer hallucination [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders] SAE steering cuts hallucination rate from 72.63% to 14.11% (small) and 86.88% to 27.33% (large-v3), at a modest clean-speech WER cost [aparin-etal-2026-whisper-hallucination-detection-mitigation-hidden-representation-steering-sparse-autoencoders]