Video foundation models vary in physics decodability, V-JEPA strongest
measured in 1 paperPunzo et al. fit frozen-feature linear probes layer-by-layer on V-JEPA, VideoMAE, and LTX-Video, evaluated on the IntPhys2 and Minimal Video Pairs intuitive-physics benchmarks [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics] V-JEPA (ViT-H) achieves the strongest results, especially with temporal-dynamics probes, VideoMAE is competitive, and LTX-Video recovers weaker but above-chance signal [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics] Physics-relevant information is weakest in early layers and peaks at intermediate-to-late depth [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics] Frame-order shuffling substantially degrades performance, confirming the signal reflects genuine temporal/physical structure rather than static appearance [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics]