Wider networks converge to more similar representations, measured by PWCCA
measured in 1 paperMorcos et al. introduce projection-weighted CCA (PWCCA), weighting each canonical correlation direction by how much of the original representation it explains, fixing SVCCA's sensitivity to low-variance directions [morcos-etal-2018-pwcca] Applied to CIFAR-10 convnets and PTB/WikiText-2 LSTMs across many runs, wider networks qualitatively converge to more similar solutions across random seeds (no coefficient is attached to the width relationship) [morcos-etal-2018-pwcca] Test accuracy is strongly anti-correlated with pairwise CCA distance (-0.96), so networks closer to the convergent solution generalize better [morcos-etal-2018-pwcca] Generalizing networks converge to more similar solutions, but memorizing networks do NOT self-cluster: they are as similar to each other as they are to a generalizing network [morcos-etal-2018-pwcca]