Trained recurrent (RNN/GRU/Mamba) policies' hidden states spontaneously converge to a stable, low-dimensional closed-orbit limit cycle that is contractive under finite-time Lyapunov analysis, and selectively randomizing only the top CCA-aligned dimensions between neural state and behavior causally collapses navigation performance back to baseline while randomizing all other dimensions preserves it
measured in 1 paperLi & Zhan (2026) train recurrent policies (RNN/GRU/Mamba, via evolution strategies or PPO) on partially-observable navigation and Procgen tasks (Jumper, Bigfish, Bossfight). Hidden states converge to a stable low-dimensional closed orbit in PCA projection, persisting across episodes and recovering after perturbation; finite-time Lyapunov-index analysis confirms optimized policies are contractive (lambda_FTLI<0) versus chaotic random networks. CCA between neural hidden states and a "Behavioral Potential Field" embedding of physical trajectories shows high canonical correlations (>0.7 sustained across 10+ modes for trained agents vs. rapid collapse after 3 dimensions for randomized controls; e.g. Jumper-PPO-GRU Pearson R=0.7473). Selectively randomizing the top CCA dimensions of an optimal hidden state collapses behavior back to baseline ("amnesia," peak convergence time reverting from ~50 to ~250 steps), while randomizing all dimensions except the top CCA ones preserves fast convergence (~50 steps) -- proving the discovered CCA subspace is both necessary and sufficient for the learned navigational behavior.