Real Atari world models IRIS and DIAMOND fall into architecture-dependent linear vs. nonlinear decodability regimes
measured in 1 paperZhang (2026) probes two real RL world models trained on real Atari Breakout/Pong replay data -- IRIS (discrete VQ-VAE-token transformer) and DIAMOND (continuous diffusion UNet) -- for game-state variables using both linear and nonlinear (MLP) probes at every layer [zhang-2026-what-do-world-models-learn-in-rl-probing-latent-representations] IRIS is approximately linearly decodable throughout (selectivity gap Delta <= 0.06 at every layer), while DIAMOND shows an inverted-V: linear decodability peaks only at the UNet's compression bottleneck and requires a nonlinear probe elsewhere -- an architecture-dependent split, not a change of probing basis [zhang-2026-what-do-world-models-learn-in-rl-probing-latent-representations] Activation patching along the probe-derived direction shifts predicted next-state token distributions consistent with the intervention (r=0.97), and multi-baseline token ablation on IRIS's spatial token grid identifies score/brick-region tokens as most causally important (rank consistency rho > 0.92) [zhang-2026-what-do-world-models-learn-in-rl-probing-latent-representations]