Real trained DDPMs lose spectral-gap dimension smoothly as training-set size shrinks into a memorization regime
measured in 1 paperAchilli, Ventura, Silvestri, Pham, Raya, Krotov, Lucibello & Ambrogioni retarget the score-Jacobian spectral-gap estimator from Ventura et al. (2025) -- built for the true (population) score -- to the empirical score of real trained DDPMs, finding new spectral gaps appear that only a theory of the empirical score, not the true score, predicts [achilli-etal-2024-losing-dimensions-geometric-memorization-diffusion] Across real DDPMs (PixelCNN++-backbone U-Nets, 24.5M-61.7M parameters) trained from scratch on MNIST, Fashion-MNIST, CIFAR-10, CelebA-HQ and LSUN-Church at 38 dataset-size splits per dataset, the spectral-gap-estimated local intrinsic dimension declines smoothly across a training-set-size window of roughly 1000-10000 samples (dataset-specific critical points A/B, e.g. CIFAR-10 A=2000/B=16000 of 50000 total) before collapsing toward zero below it [achilli-etal-2024-losing-dimensions-geometric-memorization-diffusion] A secondary participation-ratio diagnostic (effective active-sample count) tracks the same three empirical regimes -- generalization, geometric memorization, and full memorization -- as training-set size shrinks [achilli-etal-2024-losing-dimensions-geometric-memorization-diffusion]