MATH · IN · MODELS

Locally-computable ordinals form SAE-tiled 1D manifolds; numeric magnitude does not

measured in 1 paper

Bassi & Tomar test whether the curved-1D-manifold-plus-attention-twist motif generalizes across four ordinal tasks and Gemma-2-2B/9B and Qwen3-4B [bassi-tomar-2026-geometry-of-ordinal-representations] Locally-computable variables (bracket-nesting depth, table column) concentrate over 90% of per-value-centroid variance in PC1, with Gemma Scope SAE features tiling the manifold like place cells [bassi-tomar-2026-geometry-of-ordinal-representations] Indentation needs 2-3 components and numeric magnitude needs 4-5 (PC1 only 40-62%) with zero monotonic SAE features, so magnitude is strongly linearly decodable (R^2>0.93) yet forms no coherent manifold [bassi-tomar-2026-geometry-of-ordinal-representations] Qwen3-4B shows 4-10x stronger attention-head twist than Gemma, but its indentation twisters preserve ordinal rank (Spearman >=0.77) while its numeric twisters do not (<=0.23) [bassi-tomar-2026-geometry-of-ordinal-representations] Activation patching of the top-3 manifold directions drops probe accuracy over 28x more than random directions, confirming the manifolds are informationally necessary, and the character-counting motif replicates the earlier Gurnee et al. finding [gurnee-etal-2025]

Context

place-cell feature tiling (SAE), locally-computable vs. cross-position-integrated ordinal variables, attention-head twist score / alignment score / ordinality score, cross-architecture comparison (Gemma-2 vs. Qwen3), subspace ablation vs. random-direction control, numeric magnitude as a negative case (decodable but no manifold)

Papers

Geometry of Ordinal Representations in Language Models — Bassi, Saksham, Tomar, Sharvi2026 · arXiv:2607.04167
When Models Manipulate Manifolds: The Geometry of a Counting Task — Gurnee, W., Ameisen, E., Kauvar, I., Tarng, J., Pearce, A., Olah, C., Batson, J.2025 · arXiv:2601.04480