MATH · IN · MODELS

Whitening removes anisotropy and improves sentence similarity

measured in 1 paper

- Whitening, a parameter-free linear map (mean-center, then rotate and scale by the covariance eigendecomposition), makes sentence embeddings isotropic and improves cosine STS. [huang-etal-2021] - Average STS Spearman rises 62.97 to 67.76 (BERT), 59.53 to 67.72 (RoBERTa) and 64.12 to 68.30 (DistilBERT), but only 71.56 to 71.71 for the already-isotropic LaBSE. [huang-etal-2021] - Averaging the first and last layers (L1+L12) beats any single layer, and token-averaging beats the [CLS] vector by a large margin. [huang-etal-2021] - Main models are BERT-base, RoBERTa-base, DistilBERT and LaBSE, with the effect confirmed across 24 pretrained encoders in the appendix. [huang-etal-2021]

Structure

Context

covariance whitening, eigendecomposition, layer combination, token pooling

Papers

WhiteningBERT: An Easy Unsupervised Sentence Embedding Approach — Huang, Junjie, Tang, Duyu, Zhong, Wanjun, Lu, Shuai, Shou, Linjun, Gong, Ming, Jiang, Daxin, Duan, Nan2021 · arXiv:2104.01767