A persistent circle recurs in the convolutional filter weights of real trained CNNs across depths, datasets, and training runs, and fixing an idealized circle as the first layer causally improves cross-dataset generalization
measured in 1 paperPersistent homology (Vietoris-Rips barcodes) on the point cloud of convolutional filter weights confirmed a persistent one-dimensional loop (a circle) recurring at nearly all depths (layers 1-13) across roughly 1,000 real trained CNNs on MNIST, CIFAR-10, SVHN, and pretrained VGG16/VGG19 on ImageNet [carlsson-bruel-gabrielsson-2018-topological-approaches-to-deep-learning][bruel-gabrielsson-carlsson-2019-exposition-and-interpretation-of-the-topology-of-neural-networks] Weaker two- or three-circle configurations also recur in some layers and datasets; both papers loosely call this motif a "Klein bottle" only as an inherited label from Carlsson et al. (2008)'s natural-image-patch study, without independently computing non-orientability (no H2 or orientability check on the weights themselves) [carlsson-bruel-gabrielsson-2018-topological-approaches-to-deep-learning][bruel-gabrielsson-carlsson-2019-exposition-and-interpretation-of-the-topology-of-neural-networks] Fixing the first convolutional layer to an idealized discretized circle, versus a random-Gaussian control and a normally-trained control, causally raised MNIST-to-SVHN transfer accuracy from 11-12% to 28% [bruel-gabrielsson-carlsson-2019-exposition-and-interpretation-of-the-topology-of-neural-networks] Appending idealized-circle-derived features to raw pixel input sped up training by a factor of 2 on MNIST and 3.5 on SVHN, and improved MNIST-to-SVHN transfer from 10% to 22% [carlsson-bruel-gabrielsson-2018-topological-approaches-to-deep-learning] Topological simplicity of the first-layer weight point cloud (persistence of its dominant loop) correlates with the trained network's test-set generalization accuracy [bruel-gabrielsson-carlsson-2019-exposition-and-interpretation-of-the-topology-of-neural-networks]