Multi-response effective rank detects hallucination at AUROC 0.84-0.86
measured in 1 paperWang et al. compute the effective rank (Roy-Vetterli spectral-entropy dimensionality) of a matrix of embeddings sampled across multiple generated responses and layers of Llama-2-7b-chat, Llama-2-13b-chat, and Mistral-7B-v0.1 [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] Used as a hallucination signal it reaches AUROC around 0.84-0.86 across QA benchmarks, competitive with or exceeding semantic-entropy and self-consistency baselines [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] The 13B BioASQ configuration scores 0.8234, slightly below the headline 0.84-0.86 range [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty] Ablations vary the number of generations and the layer-selection strategy (middle layer vs last-5 layers) [wang-etal-2025-revisiting-hallucination-detection-with-effective-rank-based-uncertainty]