MATH · IN · MODELS

DQN and PPO agents trained on the same real navigation task develop representations invariant to different symmetry classes, depending on learning objective

measured in 1 paper

Real DQN (value-based) and PPO (policy-gradient) agents are trained on the same real navigation tasks and analyzed via MDP-reduction theory to test which symmetry classes their learned representations are invariant to [halvagal-lee-chung-2026-task-induced-representational-invariances] DQN's learned representation is invariant to MDP-homomorphism symmetries (behaviorally equivalent states under the task's reward/transition structure), while PPO's learned representation is instead invariant to action symmetries -- a discovered, algorithm-dependent difference in which symmetry group the representation collapses onto, not an architecturally imposed one [halvagal-lee-chung-2026-task-induced-representational-invariances] This objective-dependent invariance-class difference is consistent across multiple domains and has measured downstream consequences for transfer learning [halvagal-lee-chung-2026-task-induced-representational-invariances]

Context

MDP homomorphism, action symmetry, representational invariance, value-based vs. policy-gradient RL

Method

Papers

Task-Induced Representational Invariances Depend on Learning Objective in Deep RL — Halvagal, Manu Srinath, Lee, Sebastian, Chung, SueYeon2026 · arXiv:2606.01868