MATH · IN · MODELS

A Sokoban RNN's plan representation predicts 50 steps ahead and generalizes OOD

measured in 1 paper

Taufeeque et al. recover a causal plan representation from a DRC ConvLSTM Sokoban agent via logistic-regression probes, with causal validation that probed directions drive behavior [taufeeque-etal-2024-planning-in-a-recurrent-neural-network-that-plays-sokoban] The plan representation predicts the agent's future actions roughly 50 steps ahead [taufeeque-etal-2024-planning-in-a-recurrent-neural-network-that-plays-sokoban] Plan length and quality increase over the network's early internal computation steps [taufeeque-etal-2024-planning-in-a-recurrent-neural-network-that-plays-sokoban] The representation generalizes robustly to out-of-distribution puzzles far larger than any seen in training, and also explains a level-start "pacing" behavior the training incentivizes [taufeeque-etal-2024-planning-in-a-recurrent-neural-network-that-plays-sokoban]

Context

planning, world-model

Papers

Planning in a Recurrent Neural Network That Plays Sokoban — Taufeeque, Mohammad, Quirke, Philip, Li, Maximilian, Cundy, Chris, Tucker, Aaron David, Gleave, Adam, Garriga-Alonso, Adrià2024 · arXiv:2407.15421