MATH · IN · MODELS

Video world models encode motion direction as a circular population code

measured in 1 paper

Joseph et al. analyze frozen pretrained V-JEPA 2 and VideoMAE-v2 video encoders and find motion direction carried by direction-selective MLP units whose sinusoidal tuning curves tile the full angular range at an intermediate-depth Physics Emergence Zone [joseph-etal-2026-interpreting-physics-video-world-models] A sawtooth pattern in probe accuracy under successive feature orthogonalization is consistent with paired sine/cosine encodings, a circular population code rather than a single linear direction (linear-probe R^2=0.97) [joseph-etal-2026-interpreting-physics-video-world-models] Scalar physical quantities (speed, acceleration) are linearly decodable from early layers onward, a simpler shape than the direction variable's population code [joseph-etal-2026-interpreting-physics-video-world-models] Local-attention suppression at the emergence zone drops direction-decoding R^2 (0.97->0.14) and intuitive-physics accuracy (78.3%->61.7%) while leaving ImageNet nearly unchanged (33.7%->33.1%), a clean double dissociation [joseph-etal-2026-interpreting-physics-video-world-models] Steering the direction variable required jointly manipulating dozens of orthogonal probe dimensions rather than a single vector, consistent with population-code geometry [joseph-etal-2026-interpreting-physics-video-world-models]

Context

circular population code, Physics Emergence Zone, intuitive physics probing, double dissociation via targeted ablation

Papers

Interpreting Physics in Video World Models — Joseph, Sonia, Garrido, Quentin, Balestriero, Randall, Kowal, Matthew, Fel, Thomas, Bakhtiari, Shahab, Richards, Blake, Rabbat, Mike2026 · arXiv:2602.07050