An in-context-RL transformer linearly decodes XY position by layer 2
measured in 1 paperFang & Rajan train a 3-layer 512-dim causal Transformer from scratch via decision-pretraining meta-RL on gridworld (5x5) and tree-maze tasks [fang-rajan-2026-from-memories-to-maps] XY position is reliably linearly decodable from layer-2 node representations, and kernel alignment to the latent environment structure grows with in-context length, strongest at layer 2 [fang-rajan-2026-from-memories-to-maps] Representations align across differently-cued gridworld environments via cross-context kernel alignment, a cognitive-map structure emerging from in-context experience without ground-truth coordinates [fang-rajan-2026-from-memories-to-maps] Memory-token attention ablation shows the model causally selects shortcut paths in over 60% of held-out test environments [fang-rajan-2026-from-memories-to-maps]