Token-embedding orientation is shared within families but drops across them
measured in 1 paperLee et al. measure token-embedding geometry across the GPT-2, Llama-3, and Gemma-2 model families [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] Relative orientation/cosine structure is near-identical within a model family but drops sharply across families trained on different data [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] A k-NN-neighborhood PCA intrinsic-dimension estimator finds low-ID tokens form semantically coherent clusters while high-ID tokens do not [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings] EMB2EMB, a linear least-squares (OLS) map fit over 100k shared tokens, transfers CAA-style steering vectors (refusal, sycophancy, corrigibility) between differently-sized models [lee-etal-2025-shared-global-and-local-geometry-of-language-model-embeddings]