MATH · IN · MODELS

SVCCA reveals low effective dimensionality and bottom-up convergence

measured in 1 paper

Raghu et al. introduce SVCCA (SVD truncation followed by CCA) and apply it to trained CNNs (a ResNet and a plain convnet on CIFAR-10/ImageNet) and RNNs [raghu-etal-2017-svcca] Individual layers are usable at far lower effective dimension than their nominal unit count, evidencing substantial over-parameterization [raghu-etal-2017-svcca] Representations converge bottom-up: earlier layers stabilize earlier during training than later layers [raghu-etal-2017-svcca] This is the foundational CCA-based similarity method later extended by Voita et al. and contrasted against by CKA [raghu-etal-2017-svcca]

Context

SVCCA affine-invariant similarity, effective dimensionality, bottom-up training convergence

Papers

SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability — Raghu, Maithra, Gilmer, Justin, Yosinski, Jason, Sohl-Dickstein, Jascha2017 · arXiv:1706.05806