DQN and PPO agents trained on the same real navigation task develop representations invariant to different symmetry classes, depending on learning objective
measured in 1 paperReal DQN (value-based) and PPO (policy-gradient) agents are trained on the same real navigation tasks and analyzed via MDP-reduction theory to test which symmetry classes their learned representations are invariant to [halvagal-lee-chung-2026-task-induced-representational-invariances] DQN's learned representation is invariant to MDP-homomorphism symmetries (behaviorally equivalent states under the task's reward/transition structure), while PPO's learned representation is instead invariant to action symmetries -- a discovered, algorithm-dependent difference in which symmetry group the representation collapses onto, not an architecturally imposed one [halvagal-lee-chung-2026-task-induced-representational-invariances] This objective-dependent invariance-class difference is consistent across multiple domains and has measured downstream consequences for transfer learning [halvagal-lee-chung-2026-task-induced-representational-invariances]