A physics-plausibility CAV steers a video world model
measured in 1 paperAlam fits per-layer linear probes on frozen VideoMAE-base patch tokens to classify IntPhys videos as physically possible or impossible, peaking at layer 5 (70.1%) [alam-2026-causal-physics-steering-video-world-models-concept-activation-vectors] Injecting the L2-normalized layer-5 probe vector produces a clean dose-response: at alpha=+5 judgments flip to "impossible" (P=1.000), at alpha=-5 to "possible" [alam-2026-causal-physics-steering-video-world-models-concept-activation-vectors] The effect is sharply layer-localized: injection at layers 0-5 flips 25% of judgments while layers 6-11 flip 0% [alam-2026-causal-physics-steering-video-world-models-concept-activation-vectors] The physics CAV is measurably orthogonal to a motion-direction CAV (90.0 degrees) and to a random vector (85.0 degrees) [alam-2026-causal-physics-steering-video-world-models-concept-activation-vectors] Per-block physics-principle CAVs are themselves substantially non-orthogonal (O1-O2 75.7, O2-O3 86.1 degrees), a quantified structure among related concept directions [alam-2026-causal-physics-steering-video-world-models-concept-activation-vectors]