Modular-arithmetic MLPs form square-wave weights that alone generalize
measured in 1 paperSwaroop shows ReLU MLPs trained on modular arithmetic (one hidden layer, width 256, modulus 97) develop near-binary square-wave input weights, with intermediate values only near sign-change boundaries [swaroop-2026-latent-algorithmic-structure-precedes-grokking] Output weights have dominant Fourier phases satisfying phi_out = phi_a + phi_b, with frequency and phase extracted per-neuron via a discrete Fourier transform [swaroop-2026-latent-algorithmic-structure-precedes-grokking] An idealized MLP built purely from these extracted square-wave/cosine components reaches 95.5% accuracy even when derived from a model that itself scores only 0.23% [swaroop-2026-latent-algorithmic-structure-precedes-grokking] This is direct evidence that the Fourier/geometric weight structure is the causally load-bearing algorithmic content, present before the model generalizes [swaroop-2026-latent-algorithmic-structure-precedes-grokking]