MATH · IN · MODELS

Torus

OBJECTgeometric-objectsubsetK:zeromanifoldintermediatehow it's classified →

Closed manifold T² = S¹ × S¹ — the product of two independent circles. Distinguished from a bare circle by first homology: H₁(T²) = ℤ² (two independent cycles) versus H₁(S¹) = ℤ (one).

Replicationcomputed from the corpus — never hand-assigned
9 papers · no shared authors2 architecture classes · across papers3 domains · across papers8 model families · across papers
Filled = two or more values reported by papers that share no author — replication. Outlined = two or more values, but all from a single study — breadth, not replication. Grey = a single value. Derived from paper authorship and each model's architecture class, domain and family; it updates itself when a paper is added.

Definition

The torus T2=S1×S1T^2 = S^1 \times S^1 is the Cartesian product of two circles, parametrised by a pair of angles (θ1,θ2)[0,2π)2(\theta_1,\theta_2) \in [0,2\pi)^2, each taken modulo 2π2\pi independently. A standard embedding in R3\mathbb{R}^3 (radii R>r>0R > r > 0):

x=(R+rcosθ2)cosθ1,y=(R+rcosθ2)sinθ1,z=rsinθ2x = (R + r\cos\theta_2)\cos\theta_1, \quad y = (R+r\cos\theta_2)\sin\theta_1, \quad z = r\sin\theta_2

More generally, in a dd-dimensional space, T2T^2 can appear as two orthogonal 2D circular subspaces occupied independently: any pair (θ1,θ2)(\theta_1,\theta_2) that occurs on one circle can co-occur with any angle on the other.

Intuition

The surface of a donut. Walking along one direction (around the tube) returns you to your start after wrapping once; walking along the other direction (around the hole) returns you after wrapping the other way — and the two walks don’t interfere with each other. Two clock faces running independently, glued into one joint state space.

Properties

  • H1(T2)=Z2H_1(T^2) = \mathbb{Z}^2. Two independent, non-contractible loops (go around one way; go around the other way), neither reducible to a multiple of the other. This is the discriminating invariant against a bare circle (H1=ZH_1 = \mathbb{Z}, one loop): the Künneth formula gives H1(S1×S1)(H1(S1)H0(S1))(H0(S1)H1(S1))Z2H_1(S^1\times S^1) \cong \big(H_1(S^1)\otimes H_0(S^1)\big) \oplus \big(H_0(S^1)\otimes H_1(S^1)\big) \cong \mathbb{Z}^2.
  • Flat (zero Gaussian curvature) intrinsically, despite looking curved when embedded in R3\mathbb{R}^3 — the standard embedding’s curvature is purely extrinsic. T2T^2 is the unique closed, orientable, genus-1 surface, and χ(T2)=0\chi(T^2) = 0 (Euler characteristic), consistent with Gauss–Bonnet (T2KdA=2πχ(T2)=0\iint_{T^2} K \, dA = 2\pi \chi(T^2) = 0).
  • A strictly stronger claim than a circle. Finding one cyclic variable is S1S^1-evidence; the torus specifically requires two independent, orthogonal cyclic coordinates coexisting in the same joint space. Two separate, unrelated circles found in two separate experiments do not establish a torus — see Circle Exercise 5 for why independence must be checked directly, not inferred from two individually-confirmed circles.
  • Universal cover is R2\mathbb{R}^2. T2=R2/Z2T^2 = \mathbb{R}^2 / \mathbb{Z}^2 — the plane modulo integer translations in two independent directions, i.e. a square (or any fundamental parallelogram) with opposite edges identified.
  • Fundamental group π1(T2)=Z2\pi_1(T^2) = \mathbb{Z}^2, abelian — unlike the fundamental group of surfaces of genus 2\geq 2, which is non-abelian.

Relative to the “disc” (vector-addition) case

A rank-2 linear projection of T2T^2‘s standard 4-coordinate embedding (cosθ1,sinθ1,cosθ2,sinθ2)(cosθ1+cosθ2,sinθ1+sinθ2)(\cos\theta_1,\sin\theta_1,\cos\theta_2,\sin\theta_2) \mapsto (\cos\theta_1+\cos\theta_2,\sin\theta_1+\sin\theta_2) collapses the torus onto a filled, contractible 2D region (Betti numbers (1,0,0)(1,0,0), like a disc — no independent loop survives the projection) rather than a genuine T2T^2 ((1,2,1)(1,2,1)). Moisescu-Pareja, McCracken, Wiltzer, Létourneau, Daniels, Precup & Love (2025) prove that in single-hidden-layer networks trained on modular addition (a+b)modn(a+b) \bmod n, which of the two occurs is governed entirely by whether the two per-input phase variables in a “simple neuron” model are perfectly correlated (disc) or independent (torus) — see Theorem 4.1 in Theoretical / Analytical‘s Persistent homology (Betti number analysis) entry for the empirical validation.

Key evidence

The same 2025 paper resolves an apparent counter-example to the universality hypothesis: Zhong et al. (2023) had reported that “Clock” (trainable-attention) and “Pizza” (uniform-attention) networks trained on identical modular-addition data learn geometrically distinct circuits. Using PCA and persistent homology (Betti-number distributions) across 703 independently trained single-hidden-layer networks (MLP-Add, MLP-Concat, and both attention variants), the authors find Clock and Pizza are in fact nearly indistinguishable (MMD scores 0.018-0.024 between them) and both learn the disc/vector-addition manifold, not two different structures — while a third architecture (MLP-Concat) learns the genuine torus, later linearly projected down to the same disc/circle en route to the output logits. This is entirely a toy-model, synthetic-modular-arithmetic result (no pretrained language model is studied) — see modular-addition-torus-universality.

A second, independent line of evidence comes from biologically-inspired path-integration RNNs rather than language-model toy tasks. Xu, Gao, Zhang, Wei & Wu (2024) prove that a specific training-time mechanism (“conformal normalization,” which rescales the RNN’s input velocity so local neural-state displacement is proportional to physical displacement regardless of heading direction) forces the population-activity manifold to be a compact connected abelian Lie group — hence topologically a torus — and that conformal isometry further restricts the induced 2D lattice to be hexagonal rather than square, matching the hexagonal grid cells observed in the mammalian entorhinal cortex. Ablating the mechanism at training time eliminates the hexagonal structure. See conformal-normalization-provably-yields-toroidal-hexagonal-grid-cell-representations.

An earlier study using an ordinary (non-conformal) continuous-attractor path-integration RNN found that hexagonal gridness score and toroidal-manifold membership are not the same thing in practice: Schøyen, Pettersen, Holzhausen, Fyhn, Malthe-Sørenssen & Lepperød (2023) identify a cell cluster whose autocorrelogram point cloud forms a torus across multiple trained environments, distinct from the subset of cells with high hexagonal gridness score, and show by targeted ablation that only torus-cluster membership — not gridness score — is causally load-bearing for path-integration accuracy. See toroidal-cell-cluster-not-gridness-score-is-causally-load-bearing-for-path-integration.

Pellegrino & Chadwick (2025, NeurIPS 2025) derive a Riemannian pullback metric directly from a real task-trained continuous-time RNN’s own dynamics (rather than from a static point cloud) on a sequential working-memory task, and show the resulting hyper-torus’s Gaussian curvature is genuinely non-flat — switching sign at different points in time — while the pulled-back metric’s eigenvalues collapse toward zero as the network retrieves different stored memories, demonstrating the torus’s own intrinsic shape dynamically warps during task execution rather than remaining the fixed flat torus this node’s other evidence describes. See task-trained-rnns-state-space-manifold-including-a-hyper-torus-in-a-working-memory-task-develops-measured-non-flat-sign-changing-gaussian-curvature-and-eigenvalue-collapse-as-the-network-executes-its-task.

Foundational grid-cell evidence (pre-conformal-normalization)

Two earlier, foundational papers established the empirical phenomenon that Xu, Gao, Zhang, Wei & Wu (2024, above) later explained mechanistically. Cueva & Wei (2018) train an ordinary vanilla RNN end-to-end on a path-integration task (estimating 2D position from velocity in square and triangular arenas), with no architectural imposition of periodicity, and find hexagonal grid-cell-like periodic tuning (plus border and band cells) emerges in unit activity, measured via gridness score on spatial autocorrelograms — purely descriptive, no causal intervention. See cueva-wei-2018-hexagonal-grid-cell-like-periodic-spatial-tuning-quantified-by-gridness-score-emerges-in-a-trained-vanilla-rnn-with-no-architectural-imposition-of-periodicity. Banino et al. (2018) train an LSTM path-integration module feeding a deep RL navigation agent, similarly finding gridness-scored hexagonal and border-vector-cell-like tuning, and additionally show a causal result: ablating the highest-gridness-scoring units significantly degrades navigation performance, while ablating heterogeneous non-grid units does not. See banino-etal-2018-grid-cell-like-codes-emerge-in-a-trained-lstm-path-integration-module-supporting-vector-based-shortcuts-in-a-navigating-rl-agent-and-ablating-the-highest-gridness-units-degrades-navigation. A third foundational model, the Tolman-Eichenbaum Machine (Whittington, Muller, Mark, Chen, Barry, Burgess & Behrens 2020), trained on a broader range of spatial and non-spatial relational-memory tasks, develops the same measured hexagonal-periodicity grid units alongside band, border, and object-vector cells, plus hippocampal-like place cells whose remapping between distinct environments follows a structured, non-random pattern matching recorded biological data — extending the grid-code cluster beyond pure spatial path-integration into general relational-structure learning. See whittington-etal-2020-the-tolman-eichenbaum-machine-develops-real-hexagonal-grid-cells-band-cells-border-cells-and-remapping-place-cells-when-trained-on-spatial-and-relational-tasks.

Exercises

Base

  1. Write down explicit coordinates for two points on T2T^2 that differ only in θ1\theta_1 (same θ2\theta_2), and explain why they lie on the “same” small circle in the standard R3\mathbb{R}^3 embedding.
Solution

Take θ2=0\theta_2 = 0 fixed and compare θ1=0\theta_1 = 0 vs. θ1=π\theta_1 = \pi: points (R+r,0,0)(R+r,\,0,\,0) and ((R+r),0,0)(-(R+r),\,0,\,0). Both have the same θ2\theta_2, so both sit on the same latitude circle of radius R+rR+r around the central axis (the “big” loop); they differ only in where along that big loop they sit.

  1. Why is T1T^1 (as in, the 1-fold product S1S^1 alone) just the circle, and not something new?
Solution

The nn-torus is defined as Tn=(S1)nT^n = (S^1)^n, the product of nn copies of S1S^1. For n=1n=1 this is a product of one factor, i.e. no product at all — T1=S1T^1 = S^1 by definition. “Torus” as a genuinely new object (distinct from a circle) only starts at n=2n=2.

Middle

  1. Using the Künneth formula H1(A×B)(H1(A)H0(B))(H0(A)H1(B))H_1(A\times B) \cong (H_1(A)\otimes H_0(B)) \oplus (H_0(A)\otimes H_1(B)), confirm H1(T2)=Z2H_1(T^2) = \mathbb{Z}^2 from H0(S1)=ZH_0(S^1)=\mathbb{Z}, H1(S1)=ZH_1(S^1)=\mathbb{Z}.
Solution

Substituting A=B=S1A=B=S^1: H1(S1×S1)(ZZ)(ZZ)ZZ=Z2H_1(S^1\times S^1) \cong (\mathbb{Z}\otimes\mathbb{Z}) \oplus (\mathbb{Z}\otimes\mathbb{Z}) \cong \mathbb{Z}\oplus\mathbb{Z} = \mathbb{Z}^2. The two summands correspond to the two independent generating loops — “wind once around the first factor, stay fixed on the second” and vice versa.

  1. Suppose activations for two candidate cyclic variables are collected and each, considered alone, traces a clean circle under PCA. Describe a concrete statistical test (not just a visual one) for whether the joint configuration is a genuine T2T^2 rather than a 1-dimensional diagonal subset of it (recall Circle Exercise 5).
Solution

Compute, for a grid of angle pairs, whether all combinations (θ1,θ2)(\theta_1,\theta_2) are actually realized in the data (or approximately uniformly/independently so) rather than only pairs satisfying some functional relationship θ2=h(θ1)\theta_2 = h(\theta_1). Concretely: estimate the joint distribution of (θ^1,θ^2)(\hat\theta_1,\hat\theta_2) recovered from the two subspaces, and test whether it factors as (approximately) a product of marginals — e.g. via mutual information I(θ1;θ2)0I(\theta_1;\theta_2) \approx 0, or a chi-squared/independence test on binned angle pairs. If I(θ1;θ2)0I(\theta_1;\theta_2) \gg 0 (the angles are predictable from each other), the occupied set is lower-dimensional than the full product and is not T2T^2; only when the joint density is (close to) separable does the data support the two-independent-cycles claim that H1=Z2H_1=\mathbb{Z}^2 requires.

Pro

  1. Prove that T2=R2/Z2T^2 = \mathbb{R}^2/\mathbb{Z}^2 (the plane modulo the integer lattice) is homeomorphic to S1×S1S^1\times S^1.
Solution

Define ϕ:R2S1×S1\phi: \mathbb{R}^2 \to S^1\times S^1 by ϕ(x,y)=(e2πix,e2πiy)\phi(x,y) = (e^{2\pi i x}, e^{2\pi i y}) (identifying S1S^1 with unit complex numbers). This is continuous, surjective, and ϕ(x,y)=ϕ(x,y)\phi(x,y) = \phi(x',y') exactly when xxZx-x' \in \mathbb{Z} and yyZy-y'\in\mathbb{Z}, i.e. exactly when (x,y)(x,y) and (x,y)(x',y') represent the same class in R2/Z2\mathbb{R}^2/\mathbb{Z}^2. So ϕ\phi descends to a continuous bijection ϕˉ:R2/Z2S1×S1\bar\phi: \mathbb{R}^2/\mathbb{Z}^2 \to S^1\times S^1. Since R2/Z2\mathbb{R}^2/\mathbb{Z}^2 (with the quotient topology, realized concretely as the unit square with opposite edges identified) is compact and S1×S1S^1\times S^1 is Hausdorff, a continuous bijection from a compact space to a Hausdorff space is automatically a homeomorphism.

  1. A model’s activations for two cyclic variables live in a joint 4D space and are confirmed (by the independence test of Exercise 4) to realize a genuine product structure with H1=Z2H_1=\mathbb{Z}^2. A skeptic argues this is still not “one torus” but “two circles that happen to be independent,” and that the distinction is merely semantic. Give a precise mathematical sense in which they are wrong — i.e. a property the joint configuration has that the disjoint pair does not.
Solution

“Two independent circles” and “one torus” describe the same topological space when independence genuinely holds — S1×S1S^1\times S^1 is T2T^2, so the skeptic’s objection dissolves once independence is established; there is no further fact to check. But the meaningful mathematical distinction the skeptic may be gesturing at is between a disjoint union S1S1S^1 \sqcup S^1 (two separate circles, not interacting, living in the same ambient space but never combined into pairs) and a product S1×S1S^1\times S^1 (a single connected space of twice the dimension, where every point is a pair of angles). These have different homology: H1(S1S1)=Z2H_1(S^1\sqcup S^1) = \mathbb{Z}^2 and H0(S1S1)=Z2H_0(S^1\sqcup S^1)=\mathbb{Z}^2 (two connected components), whereas H1(T2)=Z2H_1(T^2)=\mathbb{Z}^2 but H0(T2)=ZH_0(T^2)=\mathbb{Z} (one connected component) — the torus is connected, the disjoint pair is not. Checking H0H_0 (number of connected components of the actual occupied point set) alongside H1H_1 is exactly the extra piece of evidence that separates “two independent circles glued into one joint object” from “two circles that never combine.”

Found in (7 observations · 8 families)

Banino et al. (2018) grid-cell navigation agent

Vector-based Navigation Using Grid-like Representations in Artificial Agents (2018)measured

Grid-like codes emerge in an LSTM path-integrator and are causally load-bearing

Details

Banino et al. train an LSTM path-integration network on foraging trajectories, then feed its representation to an A3C deep RL navigation agent [banino-etal-2018-vector-based-navigation] Gridness-score analysis of spatial autocorrelograms confirms hexagonal grid-cell-like and border-vector-cell-like periodic tuning emerges purely from training [banino-etal-2018-vector-based-navigation] The resulting agent exhibits vector-based navigation and shortcut-taking behavior [banino-etal-2018-vector-based-navigation] Ablating the highest-gridness units significantly degrades navigation while ablating non-grid units does not (effect sizes corroborated via code documentation, not the paywalled Nature text) [banino-etal-2018-vector-based-navigation]

models: LSTM path-integration module + A3C navigation agent · method: Gridness-score analysis, Causal interventions (steering)

Conformal-Normalization Path-Integration RNN

Emergence of Grid-like Representations by Training Recurrent Networks with Conformal Normalization (2024)measured

Conformal normalization provably yields a torus, only favoring hexagonal grids

Details

Xu et al. train path-integration RNNs with conformal normalization, rescaling input velocity so local neural-state displacement is direction-independently proportional to physical displacement [xu-etal-2024-conformal-grid-cells] They prove the transformations form a representation of the abelian Lie group (R^2,+), and since firing rates are bounded the population manifold is a compact connected abelian group, hence topologically a torus [xu-etal-2024-conformal-grid-cells] Hexagonality is only favored, not proven: a Fourier packing argument shows the hexagonal lattice fits the kernel better than the square, and hexagonal grids emerge numerically (gridness 0.86-0.87) [xu-etal-2024-conformal-grid-cells] Conformal normalization is a train-time constraint the geometry depends on: a train-time ablation (training without it) yields non-hexagon or stripe-like patterns instead, so the torus/hexagon structure is imposed rather than freely emergent [xu-etal-2024-conformal-grid-cells]

models: Conformal-normalization linear path-integration RNN, Conformal-normalization nonlinear (CANN-like) path-integration RNN · method: Analytical derivation, Causal interventions (steering)

Cueva & Wei (2018) path-integration RNN

Emergence of Grid-like Representations by Training Recurrent Neural Networks to Perform Spatial Localization (2018)measured

Rectangular grid-like spatial tuning emerges qualitatively in a vanilla RNN

Details

Cueva & Wei train a vanilla tanh RNN end-to-end on 2D path integration (estimating position from velocity in square and triangular arenas), with no architectural imposition of periodicity [cueva-wei-2018-grid-like-representations] Grid-like periodic spatial tuning, plus border and band cells, emerges in individual unit activity, a genuinely discovered structure [cueva-wei-2018-grid-like-representations] The identification is qualitative and visual (inspection of firing-rate maps): the paper computes no gridness score or spatial autocorrelogram [cueva-wei-2018-grid-like-representations] Grid geometry conforms to the enclosure shape: a square arena gives rectangular (not hexagonal) grids, while hexagonal/triangular arenas give closer-to-triangular grids [cueva-wei-2018-grid-like-representations]

models: Vanilla RNN (path-integration / spatial-localization task) · method: Geometric analysis

Grokking Modular-Arithmetic Transformer

Progress Measures for Grokking via Mechanistic Interpretability (2023), The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks (2023), On the Geometry and Topology of Representations: The Manifolds of Modular Addition (2025)measured

Modular-addition networks universally learn a torus projecting to a disc

Details

Nanda et al. show a 1-layer transformer trained on (a+b) mod 113 concentrates its embedding norm on 5 key Fourier frequencies, each embedding inputs as rotations that combine via trig identities into addition on the circle [nanda-etal-2023-grokking] Ablating all but the 5 key frequencies improves loss by 70% while ablating non-key frequencies does nothing, and Fourier-derived progress measures reveal grokking as three overlapping phases (memorization, circuit formation, cleanup) [nanda-etal-2023-grokking] Zhong et al. show Clock is one point in a wider phase space: networks split between the Clock circuit (multiplicative, needs attention) and a new Pizza circuit (absolute-value, a plain ReLU MLP), classified by gradient symmetricity and distance irrelevance [zhong-etal-2023] Circle-isolation interventions further reveal parallel Pizza ensembles and antipodal-pair-compensating mechanisms [zhong-etal-2023] Moisescu-Pareja et al. prove (Theorem 4.1) the first-layer representation is a torus T^2 when two phase variables are independent, or a rank-2 disc when perfectly correlated, with the disc always a linear projection of the torus [moisescu-pareja-etal-2025] Using PCA and persistent homology across 703 toy networks, Clock, Pizza, and MLP-Add are nearly indistinguishable and all learn the disc/vector-addition manifold, restoring the universality hypothesis [moisescu-pareja-etal-2025] Only MLP-Concat learns the genuine torus at layer 1, which later layers project to the same disc, so different architectures encode the same algorithm at different compression [moisescu-pareja-etal-2025]

models: Grokking Modular-Arithmetic Transformer (1 layer, 4 heads, d_model=128, mod 113) · method: PCA, Persistent homology (Betti number analysis), Fourier analysis of weights and activations, Causal interventions (steering), Algorithmic phase diagnostics (gradient symmetricity, distance irrelevance)

Clock/Pizza Modular-Arithmetic Transformer

Progress Measures for Grokking via Mechanistic Interpretability (2023), The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks (2023), On the Geometry and Topology of Representations: The Manifolds of Modular Addition (2025)measured

Modular-addition networks universally learn a torus projecting to a disc

Details

Nanda et al. show a 1-layer transformer trained on (a+b) mod 113 concentrates its embedding norm on 5 key Fourier frequencies, each embedding inputs as rotations that combine via trig identities into addition on the circle [nanda-etal-2023-grokking] Ablating all but the 5 key frequencies improves loss by 70% while ablating non-key frequencies does nothing, and Fourier-derived progress measures reveal grokking as three overlapping phases (memorization, circuit formation, cleanup) [nanda-etal-2023-grokking] Zhong et al. show Clock is one point in a wider phase space: networks split between the Clock circuit (multiplicative, needs attention) and a new Pizza circuit (absolute-value, a plain ReLU MLP), classified by gradient symmetricity and distance irrelevance [zhong-etal-2023] Circle-isolation interventions further reveal parallel Pizza ensembles and antipodal-pair-compensating mechanisms [zhong-etal-2023] Moisescu-Pareja et al. prove (Theorem 4.1) the first-layer representation is a torus T^2 when two phase variables are independent, or a rank-2 disc when perfectly correlated, with the disc always a linear projection of the torus [moisescu-pareja-etal-2025] Using PCA and persistent homology across 703 toy networks, Clock, Pizza, and MLP-Add are nearly indistinguishable and all learn the disc/vector-addition manifold, restoring the universality hypothesis [moisescu-pareja-etal-2025] Only MLP-Concat learns the genuine torus at layer 1, which later layers project to the same disc, so different architectures encode the same algorithm at different compression [moisescu-pareja-etal-2025]

models: Clock/Pizza Transformer, Model A (1 layer, constant attention alpha=0, width 128, mod 59), Clock/Pizza Transformer, Model B (1 layer, normal attention alpha=1, width 128, mod 59) · method: PCA, Persistent homology (Betti number analysis), Fourier analysis of weights and activations, Causal interventions (steering), Algorithmic phase diagnostics (gradient symmetricity, distance irrelevance)

Pellegrino & Chadwick (2025) Task-Trained Continuous-Time RNNs

RNNs Perform Task Computations by Dynamically Warping Neural Representations (2025)measured

Task-trained RNN manifolds warp with sign-changing curvature and eigenvalue collapse

Details

Pellegrino & Chadwick derive a Riemannian pullback metric on the state-space manifolds of two task-trained RNNs: a contextual evidence-integration network and a sequential working-memory network whose activity forms a hyper-torus [pellegrino-chadwick-2025-rnn-dynamic-warping] The manifold warps so as to compress irrelevant input information [pellegrino-chadwick-2025-rnn-dynamic-warping] In the working-memory network, the torus's Gaussian curvature is non-flat with both positive and negative curvature that varies spatially over the torus (not over time) [pellegrino-chadwick-2025-rnn-dynamic-warping] In the context-integration network, two metric eigenvalues (the time and irrelevant-input components) decay to zero by decision time, so that manifold becomes effectively one-dimensional [pellegrino-chadwick-2025-rnn-dynamic-warping] The study is purely geometric measurement, with no ablation of the warping mechanism [pellegrino-chadwick-2025-rnn-dynamic-warping]

models: Continuous-time rate RNN (contextual evidence-integration task, SDE-trained), Continuous-time rate RNN (sequential circular-manifold working-memory task, SDE-trained) · method: Riemannian pullback-metric curvature analysis

CARNN (Continuous Attractor RNN, Sorscher-style path-integration model)

Coherently Remapping Toroidal Cells But Not Grid Cells are Responsible for Path Integration in Virtual Agents (2023)measured

Torus-cluster membership, not gridness, is causally load-bearing for path integration

Details

Schoyen et al. train a continuous-attractor path-integration RNN (4096 units) across up to 50 environments and cluster cells by autocorrelogram similarity via UMAP+DBSCAN [schoyen-etal-2023-toroidal-cells-path-integration] A 315-cell cluster whose point cloud is torus-shaped is distinct from the high-gridness-score subset; the two only partially overlap [schoyen-etal-2023-toroidal-cells-path-integration] The torus cluster remaps coherently across environments (shared spacing and orientation) but with a consistent nonzero phase shift, mirroring biological grid-module remapping [schoyen-etal-2023-toroidal-cells-path-integration] Pruning the 315 torus-cluster cells drives path-integration decoding error to the untrained baseline, while pruning equally many high-gridness or random cells barely matters, so torus membership rather than gridness carries the computation [schoyen-etal-2023-toroidal-cells-path-integration]

models: CARNN trained on path integration across multiple (1-50) environments · method: UMAP, Geometric analysis, Causal interventions (steering)

Tolman-Eichenbaum Machine (TEM)

The Tolman-Eichenbaum Machine: Unifying Space and Relational Memory through Generalization in the Hippocampal Formation (2020)measured

The Tolman-Eichenbaum Machine develops hexagonal grid and remapping place cells

Details

Whittington et al. train the Tolman-Eichenbaum Machine, a recurrent generative model, on a range of spatial and non-spatial relational-structure tasks [whittington-etal-2020-tolman-eichenbaum-machine] It develops entorhinal-like units with real hexagonal grid periodicity, plus band, border, and object-vector cells [whittington-etal-2020-tolman-eichenbaum-machine] It also develops hippocampal-like place cells whose remapping between environments follows a structured, non-random pattern matching biological data [whittington-etal-2020-tolman-eichenbaum-machine] No causal intervention is performed; the correct citation is the Cell/bioRxiv DOI, not arXiv:1910.07663 (an unrelated paper) [whittington-etal-2020-tolman-eichenbaum-machine]

models: Tolman-Eichenbaum Machine (trained on spatial + non-spatial relational structure-learning tasks) · method: Gridness-score analysis, Place-field mapping