Modular-addition networks universally learn a torus projecting to a disc
measured in 1 paperNanda et al. show a 1-layer transformer trained on (a+b) mod 113 concentrates its embedding norm on 5 key Fourier frequencies, each embedding inputs as rotations that combine via trig identities into addition on the circle [nanda-etal-2023-grokking] Ablating all but the 5 key frequencies improves loss by 70% while ablating non-key frequencies does nothing, and Fourier-derived progress measures reveal grokking as three overlapping phases (memorization, circuit formation, cleanup) [nanda-etal-2023-grokking] Zhong et al. show Clock is one point in a wider phase space: networks split between the Clock circuit (multiplicative, needs attention) and a new Pizza circuit (absolute-value, a plain ReLU MLP), classified by gradient symmetricity and distance irrelevance [zhong-etal-2023] Circle-isolation interventions further reveal parallel Pizza ensembles and antipodal-pair-compensating mechanisms [zhong-etal-2023] Moisescu-Pareja et al. prove (Theorem 4.1) the first-layer representation is a torus T^2 when two phase variables are independent, or a rank-2 disc when perfectly correlated, with the disc always a linear projection of the torus [moisescu-pareja-etal-2025] Using PCA and persistent homology across 703 toy networks, Clock, Pizza, and MLP-Add are nearly indistinguishable and all learn the disc/vector-addition manifold, restoring the universality hypothesis [moisescu-pareja-etal-2025] Only MLP-Concat learns the genuine torus at layer 1, which later layers project to the same disc, so different architectures encode the same algorithm at different compression [moisescu-pareja-etal-2025]