The local intrinsic dimension of real Llama-2 activations traces a hunchback shape across layers that predicts generation truthfulness on real QA datasets
measured in 1 paperLocal Intrinsic Dimension (LID, via the GeoMLE estimator) is computed on the per-layer hidden activations of real Llama-2-7B and Llama-2-13B, generating answers zero-shot and few-shot on four real QA datasets -- TriviaQA, HotpotQA, TydiQA-GP, and CoQA [yin-etal-2024-truthfulness-local-intrinsic-dimension] Aggregated across Llama-2-7B's 30 layers, LID follows a hunchback shape -- rising in early layers, peaking mid-network, then declining -- that closely tracks (shifted by one or two layers behind) the layer-wise truthfulness-detection AUROC, which peaks at 0.746 averaged across datasets, outperforming entropy- and classifier-based baselines [yin-etal-2024-truthfulness-local-intrinsic-dimension] A truthfulness detector trained on one QA dataset's LID profile transfers with only modest degradation to a different QA dataset (e.g. CoQA AUROC 0.763 with same-dataset neighbors versus 0.747 with TriviaQA neighbors), indicating the LID-truthfulness relationship is not dataset-specific [yin-etal-2024-truthfulness-local-intrinsic-dimension]