SVCCA reveals low effective dimensionality and bottom-up convergence
measured in 1 paperRaghu et al. introduce SVCCA (SVD truncation followed by CCA) and apply it to trained CNNs (a ResNet and a plain convnet on CIFAR-10/ImageNet) and RNNs [raghu-etal-2017-svcca] Individual layers are usable at far lower effective dimension than their nominal unit count, evidencing substantial over-parameterization [raghu-etal-2017-svcca] Representations converge bottom-up: earlier layers stabilize earlier during training than later layers [raghu-etal-2017-svcca] This is the foundational CCA-based similarity method later extended by Voita et al. and contrasted against by CKA [raghu-etal-2017-svcca]
Context
SVCCA affine-invariant similarity, effective dimensionality, bottom-up training convergence
Confirmed in models
Papers
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability — Raghu, Maithra, Gilmer, Justin, Yosinski, Jason, Sohl-Dickstein, Jascha