Lower input-sequence intrinsic dimension predicts more verbatim memorization
measured in 1 paperArnold treats each training text as a point cloud of BERT contextual embeddings and estimates its TwoNN intrinsic dimension, a property of the input sequence rather than a layer-wise hidden-state profile of the studied model [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] Across GPT-Neo 125M/1.3B/2.7B and GPT-J-6B, lower-ID sequences are more likely to be verbatim-memorized and higher-ID sequences less likely, so intrinsic dimension acts as a suppressive signal for memorization [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] In the low-duplication regime, memorization declines inversely with intrinsic dimensionality across all model sizes [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] The relationship is purely observational: no detection classifier and no mutual-information analysis are performed [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension] Scope is restricted to exact-duplicate verbatim memorization under greedy decoding on 1,000 Pile sequences of 150 tokens [arnold-2025-memorization-in-language-models-through-the-lens-of-intrinsic-dimension]