Two independently-trained JEPA world models converge to a linear isomorphism
measured in 1 paperZhang et al. train pairs of ViT-S/16 I-JEPA encoders fully independently on different views of the same scenes (smallNORB, nuScenes multi-camera, ImageNet augmentation views) [zhang-etal-2026-social-jepa] They define geometric isomorphism as an invertible linear map z2 approximately W z1, fit by closed-form ridge regression and quantified by MSE, R-squared, linear CKA, distance-structure consistency and neighborhood overlap [zhang-etal-2026-social-jepa] Best case (smallNORB) reaches MSE 0.036, R-squared 0.891, DSC 0.872, a multiply-corroborated approximate linear isometry between independently-learned spaces [zhang-etal-2026-social-jepa] They prove the JEPA loss is invariant under GL(d) reparameterization of the encoder, giving a theoretical reason to expect this convergence class [zhang-etal-2026-social-jepa] The fitted map transfers a linear probe zero-shot with no gradient steps and enables teacher-student representation migration at 0.28x the FLOPs of training from scratch [zhang-etal-2026-social-jepa]