A trained transformer forms separate subspaces for constants and variables
measured in 1 paperWu, Geiger & Milliere train a 12-layer 8-head GPT-2-style transformer from scratch (37.8M params, RoPE) on synthetic variable-assignment programs to >99.9% accuracy [wu-geiger-milliere-2025-variable-binding-symbolic-programs] PCA plus L1-regularized probing finds separate residual-stream subspaces for numerical-constant (10 components) and variable-name (26 components) information [wu-geiger-milliere-2025-variable-binding-symbolic-programs] UMAP shows increasing cluster separation across training [wu-geiger-milliere-2025-variable-binding-symbolic-programs] Interchange interventions swapping only the selected subspace between original and counterfactual programs causally validate each subspace role [wu-geiger-milliere-2025-variable-binding-symbolic-programs]