MATH · IN · MODELS

A directional-coherence score separates V-JEPA from cosine-similar rivals

measured in 1 paper

Alrasheed et al. linearly probe four frozen video foundation models (V-JEPA2.1, V-JEPA2, VideoPrism, VideoMAEv2) at matched ViT-L for push/pull action directions on Something-Something v2 [alrasheed-etal-2026-latent-video-prediction-learns-better-world-models] Their Directional Semantic Coherence Score (DSCS = r_sem x (1 - cos_rev)) is several times higher for V-JEPA than for VideoPrism and VideoMAEv2 [alrasheed-etal-2026-latent-video-prediction-learns-better-world-models] VideoPrism alone maintains representational similarity above 0.98 under severe patch dropout yet collapses to 2.7% top-1, while V-JEPA2.1 retains 46.1% [alrasheed-etal-2026-latent-video-prediction-learns-better-world-models] V-JEPA's cosine similarity under corruption is in fact lower than VideoPrism's, so DSCS captures oriented-axis structure invisible to bulk cosine similarity [alrasheed-etal-2026-latent-video-prediction-learns-better-world-models] The analysis is purely observational, with no causal intervention on the representations [alrasheed-etal-2026-latent-video-prediction-learns-better-world-models]

Context

directional semantic consistency, cosine similarity dissociation, video world models, patch-dropout robustness

Papers

Latent Video Prediction Learns Better World Models — Alrasheed, Ali J., Yazdan Parast, Aryan, Azam, Basim, Bailey, James, Akhtar, Naveed2026 · arXiv:2605.15618