MATH · IN · MODELS

Text-only LMs linearly recover CIELAB color-space structure

measured in 1 paper

Abdou et al. test whether text-only LMs encode perceptual color structure against the WCS/XKCD-derived 3D CIELAB space, with no visual grounding [abdou-etal-2021-color-perceptual-structure] RSA (Kendall's tau) between color-term embeddings and CIELAB distances is significant in the color-context configuration (BERT-large 0.24, ELECTRA 0.23, RoBERTa 0.19; the random null is not) [abdou-etal-2021-color-perceptual-structure] A lasso linear mapping onto 3D CIELAB with control-task selectivity is high for all three families (0.76-0.78), needing only ~10-40 dimensions to explain 0.4-0.7 of variance [abdou-etal-2021-color-perceptual-structure] Alignment scales with model size (BERT-mini tau 0.077 -> BERT-base 0.162) with a warm/cool recovery asymmetry; no causal intervention is performed [abdou-etal-2021-color-perceptual-structure]

Context

RSA (Kendall's tau) against CIELAB ground-truth distances, control-task-corrected linear-mapping selectivity (Hewitt & Liang 2019), model-size scaling of perceptual alignment, warm/cool color recovery asymmetry, textual surprisal correlates weakly with recoverability

Papers

Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color — Abdou, Mostafa, Kulmizev, Artur, Hershcovich, Daniel, Frank, Stella, Pavlick, Ellie, Søgaard, Anders2021 · arXiv:2109.06129