Contextualized embeddings lie on a low-dimensional manifold rising with depth
measured in 1 paperCai et al. estimate Local Intrinsic Dimension (k-NN expansion model, K=100) for BERT, DistilBERT, GPT, GPT-2, and ELMo on Penn Treebank [cai-etal-2021] Mean LID is far below ambient dimension in every model (BERT 5.6, DistilBERT 7.3, GPT 6.8, GPT-2 7.0, ELMo 9.1) versus 18.0-26.1 for static GloVe/word2vec embeddings measured the same way [cai-etal-2021] LID increases nearly linearly with layer depth across all models, a monotonic-growth trajectory unlike the expansion-then-compression profile TwoNN finds in protein/image models (a discrepancy left unreconciled) [cai-etal-2021] A qualitative PCA visualization shows GPT and GPT-2 later-layer embeddings resemble a "Swiss Roll" manifold thickening with depth, while BERT and DistilBERT do not [cai-etal-2021]