MATH · IN · MODELS

Video foundation models vary in physics decodability, V-JEPA strongest

measured in 1 paper

Punzo et al. fit frozen-feature linear probes layer-by-layer on V-JEPA, VideoMAE, and LTX-Video, evaluated on the IntPhys2 and Minimal Video Pairs intuitive-physics benchmarks [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics] V-JEPA (ViT-H) achieves the strongest results, especially with temporal-dynamics probes, VideoMAE is competitive, and LTX-Video recovers weaker but above-chance signal [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics] Physics-relevant information is weakest in early layers and peaks at intermediate-to-late depth [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics] Frame-order shuffling substantially degrades performance, confirming the signal reflects genuine temporal/physical structure rather than static appearance [punzo-etal-2026-do-video-foundation-models-understand-intuitive-physics]

Context

intuitive-physics, video-representation

Papers

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis — Punzo, Samuele, Caselli, Niccolò, Pantelidis, Ippokratis, Massafra, Francesco, Lo Sardo, Salvatore, Salehi, Mohammadreza2026 · arXiv:2606.09646