MATH · IN · MODELS

Wider networks converge to more similar representations, measured by PWCCA

measured in 1 paper

Morcos et al. introduce projection-weighted CCA (PWCCA), weighting each canonical correlation direction by how much of the original representation it explains, fixing SVCCA's sensitivity to low-variance directions [morcos-etal-2018-pwcca] Applied to CIFAR-10 convnets and PTB/WikiText-2 LSTMs across many runs, wider networks qualitatively converge to more similar solutions across random seeds (no coefficient is attached to the width relationship) [morcos-etal-2018-pwcca] Test accuracy is strongly anti-correlated with pairwise CCA distance (-0.96), so networks closer to the convergent solution generalize better [morcos-etal-2018-pwcca] Generalizing networks converge to more similar solutions, but memorizing networks do NOT self-cluster: they are as similar to each other as they are to a generalizing network [morcos-etal-2018-pwcca]

Context

projection-weighted CCA, representational similarity across width, generalization vs memorization geometry

Papers

Insights on Representational Similarity in Neural Networks with Canonical Correlation Analysis — Morcos, Ari S., Raghu, Maithra, Bengio, Samy2018 · arXiv:1806.05759