Definition
The torus is the Cartesian product of two circles, parametrised by a pair of angles , each taken modulo independently. A standard embedding in (radii ):
More generally, in a -dimensional space, can appear as two orthogonal 2D circular subspaces occupied independently: any pair that occurs on one circle can co-occur with any angle on the other.
Intuition
The surface of a donut. Walking along one direction (around the tube) returns you to your start after wrapping once; walking along the other direction (around the hole) returns you after wrapping the other way — and the two walks don’t interfere with each other. Two clock faces running independently, glued into one joint state space.
Properties
- . Two independent, non-contractible loops (go around one way; go around the other way), neither reducible to a multiple of the other. This is the discriminating invariant against a bare circle (, one loop): the Künneth formula gives .
- Flat (zero Gaussian curvature) intrinsically, despite looking curved when embedded in — the standard embedding’s curvature is purely extrinsic. is the unique closed, orientable, genus-1 surface, and (Euler characteristic), consistent with Gauss–Bonnet ().
- A strictly stronger claim than a circle. Finding one cyclic variable is -evidence; the torus specifically requires two independent, orthogonal cyclic coordinates coexisting in the same joint space. Two separate, unrelated circles found in two separate experiments do not establish a torus — see Circle Exercise 5 for why independence must be checked directly, not inferred from two individually-confirmed circles.
- Universal cover is . — the plane modulo integer translations in two independent directions, i.e. a square (or any fundamental parallelogram) with opposite edges identified.
- Fundamental group , abelian — unlike the fundamental group of surfaces of genus , which is non-abelian.
Relative to the “disc” (vector-addition) case
A rank-2 linear projection of ‘s standard 4-coordinate embedding collapses the torus onto a filled, contractible 2D region (Betti numbers , like a disc — no independent loop survives the projection) rather than a genuine (). Moisescu-Pareja, McCracken, Wiltzer, Létourneau, Daniels, Precup & Love (2025) prove that in single-hidden-layer networks trained on modular addition , which of the two occurs is governed entirely by whether the two per-input phase variables in a “simple neuron” model are perfectly correlated (disc) or independent (torus) — see Theorem 4.1 in Theoretical / Analytical‘s Persistent homology (Betti number analysis) entry for the empirical validation.
Key evidence
The same 2025 paper resolves an apparent counter-example to the universality hypothesis: Zhong et al. (2023) had reported that “Clock” (trainable-attention) and “Pizza” (uniform-attention) networks trained on identical modular-addition data learn geometrically distinct circuits. Using PCA and persistent homology (Betti-number distributions) across 703 independently trained single-hidden-layer networks (MLP-Add, MLP-Concat, and both attention variants), the authors find Clock and Pizza are in fact nearly indistinguishable (MMD scores 0.018-0.024 between them) and both learn the disc/vector-addition manifold, not two different structures — while a third architecture (MLP-Concat) learns the genuine torus, later linearly projected down to the same disc/circle en route to the output logits. This is entirely a toy-model, synthetic-modular-arithmetic result (no pretrained language model is studied) — see modular-addition-torus-universality.
A second, independent line of evidence comes from biologically-inspired path-integration RNNs rather than language-model toy tasks. Xu, Gao, Zhang, Wei & Wu (2024) prove that a specific training-time mechanism (“conformal normalization,” which rescales the RNN’s input velocity so local neural-state displacement is proportional to physical displacement regardless of heading direction) forces the population-activity manifold to be a compact connected abelian Lie group — hence topologically a torus — and that conformal isometry further restricts the induced 2D lattice to be hexagonal rather than square, matching the hexagonal grid cells observed in the mammalian entorhinal cortex. Ablating the mechanism at training time eliminates the hexagonal structure. See conformal-normalization-provably-yields-toroidal-hexagonal-grid-cell-representations.
An earlier study using an ordinary (non-conformal) continuous-attractor path-integration RNN found that hexagonal gridness score and toroidal-manifold membership are not the same thing in practice: Schøyen, Pettersen, Holzhausen, Fyhn, Malthe-Sørenssen & Lepperød (2023) identify a cell cluster whose autocorrelogram point cloud forms a torus across multiple trained environments, distinct from the subset of cells with high hexagonal gridness score, and show by targeted ablation that only torus-cluster membership — not gridness score — is causally load-bearing for path-integration accuracy. See toroidal-cell-cluster-not-gridness-score-is-causally-load-bearing-for-path-integration.
Pellegrino & Chadwick (2025, NeurIPS 2025) derive a Riemannian pullback
metric directly from a real task-trained continuous-time RNN’s own
dynamics (rather than from a static point cloud) on a sequential
working-memory task, and show the resulting hyper-torus’s Gaussian
curvature is genuinely non-flat — switching sign at different points
in time — while the pulled-back metric’s eigenvalues collapse toward
zero as the network retrieves different stored memories, demonstrating
the torus’s own intrinsic shape dynamically warps during task
execution rather than remaining the fixed flat torus this node’s other
evidence describes. See
task-trained-rnns-state-space-manifold-including-a-hyper-torus-in-a-working-memory-task-develops-measured-non-flat-sign-changing-gaussian-curvature-and-eigenvalue-collapse-as-the-network-executes-its-task.
Foundational grid-cell evidence (pre-conformal-normalization)
Two earlier, foundational papers established the empirical phenomenon
that Xu, Gao, Zhang, Wei & Wu (2024, above) later explained
mechanistically. Cueva & Wei (2018) train an ordinary vanilla RNN
end-to-end on a path-integration task (estimating 2D position from
velocity in square and triangular arenas), with no architectural
imposition of periodicity, and find hexagonal grid-cell-like periodic
tuning (plus border and band cells) emerges in unit activity, measured
via gridness score on spatial autocorrelograms — purely descriptive,
no causal intervention. See
cueva-wei-2018-hexagonal-grid-cell-like-periodic-spatial-tuning-quantified-by-gridness-score-emerges-in-a-trained-vanilla-rnn-with-no-architectural-imposition-of-periodicity.
Banino et al. (2018) train an LSTM path-integration module feeding a
deep RL navigation agent, similarly finding gridness-scored hexagonal
and border-vector-cell-like tuning, and additionally show a causal
result: ablating the highest-gridness-scoring units significantly
degrades navigation performance, while ablating heterogeneous non-grid
units does not. See
banino-etal-2018-grid-cell-like-codes-emerge-in-a-trained-lstm-path-integration-module-supporting-vector-based-shortcuts-in-a-navigating-rl-agent-and-ablating-the-highest-gridness-units-degrades-navigation.
A third foundational model, the Tolman-Eichenbaum Machine (Whittington,
Muller, Mark, Chen, Barry, Burgess & Behrens 2020), trained on a broader
range of spatial and non-spatial relational-memory tasks, develops the
same measured hexagonal-periodicity grid units alongside band, border,
and object-vector cells, plus hippocampal-like place cells whose
remapping between distinct environments follows a structured,
non-random pattern matching recorded biological data — extending the
grid-code cluster beyond pure spatial path-integration into general
relational-structure learning. See
whittington-etal-2020-the-tolman-eichenbaum-machine-develops-real-hexagonal-grid-cells-band-cells-border-cells-and-remapping-place-cells-when-trained-on-spatial-and-relational-tasks.
Exercises
Base
- Write down explicit coordinates for two points on that differ only in (same ), and explain why they lie on the “same” small circle in the standard embedding.
Solution
Take fixed and compare vs. : points and . Both have the same , so both sit on the same latitude circle of radius around the central axis (the “big” loop); they differ only in where along that big loop they sit.
- Why is (as in, the 1-fold product alone) just the circle, and not something new?
Solution
The -torus is defined as , the product of copies of . For this is a product of one factor, i.e. no product at all — by definition. “Torus” as a genuinely new object (distinct from a circle) only starts at .
Middle
- Using the Künneth formula , confirm from , .
Solution
Substituting : . The two summands correspond to the two independent generating loops — “wind once around the first factor, stay fixed on the second” and vice versa.
- Suppose activations for two candidate cyclic variables are collected and each, considered alone, traces a clean circle under PCA. Describe a concrete statistical test (not just a visual one) for whether the joint configuration is a genuine rather than a 1-dimensional diagonal subset of it (recall Circle Exercise 5).
Solution
Compute, for a grid of angle pairs, whether all combinations are actually realized in the data (or approximately uniformly/independently so) rather than only pairs satisfying some functional relationship . Concretely: estimate the joint distribution of recovered from the two subspaces, and test whether it factors as (approximately) a product of marginals — e.g. via mutual information , or a chi-squared/independence test on binned angle pairs. If (the angles are predictable from each other), the occupied set is lower-dimensional than the full product and is not ; only when the joint density is (close to) separable does the data support the two-independent-cycles claim that requires.
Pro
- Prove that (the plane modulo the integer lattice) is homeomorphic to .
Solution
Define by (identifying with unit complex numbers). This is continuous, surjective, and exactly when and , i.e. exactly when and represent the same class in . So descends to a continuous bijection . Since (with the quotient topology, realized concretely as the unit square with opposite edges identified) is compact and is Hausdorff, a continuous bijection from a compact space to a Hausdorff space is automatically a homeomorphism.
- A model’s activations for two cyclic variables live in a joint 4D space and are confirmed (by the independence test of Exercise 4) to realize a genuine product structure with . A skeptic argues this is still not “one torus” but “two circles that happen to be independent,” and that the distinction is merely semantic. Give a precise mathematical sense in which they are wrong — i.e. a property the joint configuration has that the disjoint pair does not.
Solution
“Two independent circles” and “one torus” describe the same topological space when independence genuinely holds — is , so the skeptic’s objection dissolves once independence is established; there is no further fact to check. But the meaningful mathematical distinction the skeptic may be gesturing at is between a disjoint union (two separate circles, not interacting, living in the same ambient space but never combined into pairs) and a product (a single connected space of twice the dimension, where every point is a pair of angles). These have different homology: and (two connected components), whereas but (one connected component) — the torus is connected, the disjoint pair is not. Checking (number of connected components of the actual occupied point set) alongside is exactly the extra piece of evidence that separates “two independent circles glued into one joint object” from “two circles that never combine.”