An affine map between two vision models matches their PCA components
measured in 1 paperMoayeri et al. fit an affine least-squares map between the activation spaces of independently-trained vision models (supervised/robust ResNets, Swin/DeiT/ConViT, self-supervised MoCo/DINO-ViT-S/SimCLR ResNets and ViTs, and CLIP), scored by R^2 [moayeri-etal-2023-text-to-concept-and-back-via-cross-model-alignment] The fitted map explains R^2 above 0.6, and the top principal components of the two aligned spaces correspond approximately one-to-one [moayeri-etal-2023-text-to-concept-and-back-via-cross-model-alignment] Aligning a vision encoder into CLIP's concept space enables zero-shot concept-bottleneck classification at up to 93.8% accuracy and over 92% concept-to-text relevance [moayeri-etal-2023-text-to-concept-and-back-via-cross-model-alignment] The analysis is passive, with no causal intervention on either model [moayeri-etal-2023-text-to-concept-and-back-via-cross-model-alignment]