MATH · IN · MODELS

Persistent-homology H1 cycle strength rises sharply at grokking

measured in 1 paper

Tang et al. apply persistent homology (Vietoris-Rips, degrees 0-1) to token-embedding and hidden-state point clouds from modular-addition models mod {113,149,197}, using a 2-layer transformer and a 3-hidden-layer MLP [tang-etal-2026-topological-signatures-of-grokking] H1 (loop/cycle) persistence rises sharply and reproducibly at the grokking transition (e.g. transformer p=197 max persistence ~0.075 to 0.20-0.25) [tang-etal-2026-topological-signatures-of-grokking] This is an independent topological confirmation of the circular structure Nanda et al. (2023) found via Fourier analysis, now also demonstrated in an MLP architecture [tang-etal-2026-topological-signatures-of-grokking] Local intrinsic dimension (TwoNN) collapses from ~20-25 to ~5 at the transition, alongside the topological cycle strengthening [tang-etal-2026-topological-signatures-of-grokking] Persistence statistics correlate with test accuracy (max H1 rho up to 0.81; total H0 rho down to -0.91), a correlational not causal linkage [tang-etal-2026-topological-signatures-of-grokking] A label-permutation ablation breaks the topology-generalization link above ~10-20% corruption, and an MNIST control shows no sharp topological transition [tang-etal-2026-topological-signatures-of-grokking]

Structure

Context

persistent homology (Vietoris-Rips filtration, H0/H1 persistence statistics) as a topological-invariant measurement of circular/cyclic structure, independent of Fourier/PCA-based methods, sharp, reproducible rise in H1 max/total persistence specifically at the grokking phase transition, replicated across 3 primes and 2 architectures (transformer, MLP), correlational (not causal) linkage between persistent-homology statistics and test accuracy (Spearman rho up to 0.81, and up to -0.91 for H0), local intrinsic dimension collapse (TwoNN estimator, ~20-25 to ~5) co-occurring with the topological transition, label-permutation ablation showing the topology-generalization association breaks down under high training-label corruption, dissociation from a non-grokking control task (MNIST), which shows no sharp topological transition

Papers

Topological Signatures of Grokking — Tang, Yifan, Wang, Qiquan, García-Redondo, Inés, Monod, Anthea2026 · arXiv:2605.06352