Intrinsic dimension expands then compresses; semantics peak at the minimum
measured in 1 paperValeriani et al. measure TwoNN intrinsic dimension layer-by-layer across ESM-2 protein models (35M/650M/3B) and iGPT image transformers (S/M/L), two non-textual self-supervised domains [valeriani-etal-2023] Every model follows the same arc: a sharp early expansion (ID peak ~20-32 in the first third) followed by compression to a low plateau or local minimum (ID 5-7 in ESM-2; ~22 in iGPT) [valeriani-etal-2023] Semantic content peaks at the most compressed layer: ESM-2 homology overlap is best at the plateau (a ~6% improvement over the last layer) and iGPT class-label overlap peaks at the ID minimum, scaling with model size [valeriani-etal-2023] A preliminary appendix on Llama-2-70B (SST) shows a more complex three-peak profile with sentiment overlap highest at the first local ID minimum, flagged by the authors as future work [valeriani-etal-2023] Li et al. independently confirm middle-layer compression on Gemma-2-2B SAE-feature point clouds using a covariance-eigenvalue-spectrum test against a Marchenko-Pastur null [li-etal-2024-geometry] The top-100 eigenvalues decay as a power law, steepest at layer 12 (slope -0.47) and shallower at layers 0 (-0.24) and 24 (-0.25) [li-etal-2024-geometry] A k-NN-entropy clustering-entropy measure reaches a minimum at middle layers, locating the same information bottleneck via SAE-feature rather than raw-activation geometry [li-etal-2024-geometry]