MATH · IN · MODELS

A non-monotonic spatial-frequency accessibility profile across depth in real CLIP and DINOv2 vision transformers

measured in 1 paper

Kitessa & Zhao define spectral accessibility, the ridge-regression-probe recoverability of ground-truth radial Fourier-energy bands from a frozen vision encoder's layer-wise representation, and a Residual Spectral Loss (RSL) that subtracts a dimension-matched random-projection baseline to isolate the learned transformation's own effect [kitessa-zhao-2026-quantifying-spectral-accessibility-vision-representations] Across real CLIP ViT-L/14, CLIP ViT-B/32, and DINOv2 ViT-B on ImageNet and MS-COCO images, accessibility rises roughly 2-3x from the convolutional stem (alpha approximately 0.26-0.42) to a mid-layer peak (alpha approximately 0.71-0.76), then falls 17-27 percent toward the output [kitessa-zhao-2026-quantifying-spectral-accessibility-vision-representations] CLIP's final projection layer is spectrally neutral (RSL approximately plus/minus 0.035, matching the random baseline), while DINOv2's CLS-token pooling induces a large, statistically significant spectral loss (RSL +0.10 to +0.22, confidence interval excluding zero) concentrated at low-to-mid frequencies [kitessa-zhao-2026-quantifying-spectral-accessibility-vision-representations]

Method

Papers

Beyond Compression: Quantifying Spectral Accessibility in Vision Representations — Kitessa, Akayou A., Zhao, Yijun2026 · arXiv:2606.03795