SAE directions show CLIP encodes abstract concepts, DINOv2 fragmented visual features
measured in 1 paperStevens et al. train 24K-width sparse autoencoders on frozen patch-level activations from CLIP ViT-B/16 and DINOv2 ViT-B/14 using ImageNet-1K [stevens-etal-2025-interpretable-and-testable-vision-features-via-saes] CLIP's activation geometry encodes abstract, style-invariant semantic/cultural concepts (country identity, "accident") largely absent or fragmented in DINOv2 [stevens-etal-2025-interpretable-and-testable-vision-features-via-saes] DINOv2's more visually granular geometry instead organizes around low-level visual pattern similarity [stevens-etal-2025-interpretable-and-testable-vision-features-via-saes] The discovered directions are causally responsible for predictions via patch-level activation edits, flipping ADE20K segmentation classes without retraining the ViT or task heads [stevens-etal-2025-interpretable-and-testable-vision-features-via-saes]