MATH · IN · MODELS

A diff-of-means hallucination direction causally raises hallucination dose-dependently

measured in 1 paper

Cherukuri & Varshney formalize hallucinations as basin attractors across Llama-3.2-1B/3B, Gemma-2-2B, Qwen2.5-1.5B, Llama-3.1-8B, and Mistral-7B-v0.3 [cherukuri-varshney-2026-hallucination-basins-a-dynamic-framework-for-understanding-and-controlling-llm-hallucinations] Basin separation is strongly task-dependent: factoid tasks show sharp centroid separation (variance ratio up to 4.55, AUROC up to 1.000 on MuSiQue) while summarization and TruthfulQA hover near chance [cherukuri-varshney-2026-hallucination-basins-a-dynamic-framework-for-understanding-and-controlling-llm-hallucinations] Interpolating factual hidden states toward the hallucination centroid via a diff-of-means steering vector produces a monotonic dose-response increase in hallucination probability, exceeding random and orthogonal controls [cherukuri-varshney-2026-hallucination-basins-a-dynamic-framework-for-understanding-and-controlling-llm-hallucinations]

Context

a diff-of-means steering direction between factual and hallucinated hidden-state centroids, causally validated via dose-response interpolation against random/orthogonal-direction controls, task-dependent centroid separation and variance-ratio geometry (point-attractor vs. near-chance) predicting when a hallucination-detection classifier will work at all

Papers

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations — Cherukuri, Kalyan, Varshney, Lav R.2026 · arXiv:2604.04743