MATH · IN · MODELS

A single token's residual linearly decodes an entire maze

measured in 1 paper

Ivanitskiy et al. train small (<10M) GPT transformers from scratch on synthetic maze token sequences with one-hot orthogonal input tokens [ivanitskiy-etal-2023] Linear probes on a single fixed token's residual stream reconstruct the entire maze's wall structure (>90% wall accuracy, peaking at layer 2), a striking single-site compression [ivanitskiy-etal-2023] Learned embedding vectors develop emergent spatial structure, with Manhattan distance correlating with embedding L1 distance at short range despite orthogonal inputs [ivanitskiy-etal-2023] Candidate "Adjacency Heads" attend to path-adjacent tokens in-context, correlational evidence not causally ablated in the paper [ivanitskiy-etal-2023] Sharp improvements in linear maze decodability coincide with sharp task-generalization gains, a grokking-like transition [ivanitskiy-etal-2023]

Context

world models, toy/synthetic sequence model, single-token compression, emergent embedding geometry, grokking, attention circuits

Papers

Structured World Representations in Maze-Solving Transformers — Ivanitskiy, Michael I., Spies, Alex F., Räuker, Tilman, Corlouer, Guillaume, Mathwin, Chris, Quirke, Lucia, Rager, Can, Shah, Rusheb, Valentine, Dan, Behn, Cecilia Diniz, Inoue, Katsumi, Fung, Samy Wu2023 · arXiv:2312.02566