MATH · IN · MODELS

Real trained DDPMs lose spectral-gap dimension smoothly as training-set size shrinks into a memorization regime

measured in 1 paper

Achilli, Ventura, Silvestri, Pham, Raya, Krotov, Lucibello & Ambrogioni retarget the score-Jacobian spectral-gap estimator from Ventura et al. (2025) -- built for the true (population) score -- to the empirical score of real trained DDPMs, finding new spectral gaps appear that only a theory of the empirical score, not the true score, predicts [achilli-etal-2024-losing-dimensions-geometric-memorization-diffusion] Across real DDPMs (PixelCNN++-backbone U-Nets, 24.5M-61.7M parameters) trained from scratch on MNIST, Fashion-MNIST, CIFAR-10, CelebA-HQ and LSUN-Church at 38 dataset-size splits per dataset, the spectral-gap-estimated local intrinsic dimension declines smoothly across a training-set-size window of roughly 1000-10000 samples (dataset-specific critical points A/B, e.g. CIFAR-10 A=2000/B=16000 of 50000 total) before collapsing toward zero below it [achilli-etal-2024-losing-dimensions-geometric-memorization-diffusion] A secondary participation-ratio diagnostic (effective active-sample count) tracks the same three empirical regimes -- generalization, geometric memorization, and full memorization -- as training-set size shrinks [achilli-etal-2024-losing-dimensions-geometric-memorization-diffusion]

Papers

Losing Dimensions: Geometric Memorization in Generative Diffusion — Achilli, Beatrice, Ventura, Enrico, Silvestri, Gianluigi, Pham, Bao, Raya, Gabriel, Krotov, Dmitry, Lucibello, Carlo, Ambrogioni, Luca2024 · arXiv:2410.08727