MATH · IN · MODELS

Inverting video diffusion transformers makes physical plausibility linearly decodable

measured in 1 paper

Esmati et al. approximately invert the deterministic diffusion sampling of WAN-1.3B, CogVideoX-2B, and LTX-2B to recover intermediate states and attention maps [esmati-etal-2026-the-invisible-hand-of-physics] Physical plausibility is linearly decodable from the recovered states at ~81.27% average accuracy across IntPhys and InfLevel [esmati-etal-2026-the-invisible-hand-of-physics] This exceeds V-JEPA2 ViT-L and VideoMAE-Large baselines probed on their own forward-pass activations, so diffusion models implicitly encode physical structure recoverable only via inversion [esmati-etal-2026-the-invisible-hand-of-physics]

Context

intuitive-physics, video-diffusion

Papers

The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show — Esmati, Parsa, Nath, Somjit, Hofmann, Katja, Nowrouzezahrai, Derek, Ebrahimi Kahou, Samira, Mirmehdi, Majid2026 · arXiv:2606.05328