Whitening removes anisotropy and improves sentence similarity
measured in 1 paper- Whitening, a parameter-free linear map (mean-center, then rotate and scale by the covariance eigendecomposition), makes sentence embeddings isotropic and improves cosine STS. [huang-etal-2021] - Average STS Spearman rises 62.97 to 67.76 (BERT), 59.53 to 67.72 (RoBERTa) and 64.12 to 68.30 (DistilBERT), but only 71.56 to 71.71 for the already-isotropic LaBSE. [huang-etal-2021] - Averaging the first and last layers (L1+L12) beats any single layer, and token-averaging beats the [CLS] vector by a large margin. [huang-etal-2021] - Main models are BERT-base, RoBERTa-base, DistilBERT and LaBSE, with the effect confirmed across 24 pretrained encoders in the appendix. [huang-etal-2021]
Structure
Context
covariance whitening, eigendecomposition, layer combination, token pooling
Confirmed in models
Method
Papers
WhiteningBERT: An Easy Unsupervised Sentence Embedding Approach — Huang, Junjie, Tang, Duyu, Zhong, Wanjun, Lu, Shuai, Shou, Linjun, Gong, Ming, Jiang, Daxin, Duan, Nan