Composed-subword vs whole-word spaces show family-dependent isometry
measured in 1 paperPeng, Chai & Sogaard fit an orthogonal Procrustes map between composed-subword and whole-word embedding spaces across instruction-tuned LLMs, scored by Precision@1 retrieval [peng-chai-sogaard-2025-subword-compositionality] Simple addition consistently outperforms other subword-composition operations [peng-chai-sogaard-2025-subword-compositionality] Three family patterns emerge: Aya-expanse and Gemma show high composed-vs-whole-word isometry, Llama 3/3.1 very little, others moderate but dropping late in the network [peng-chai-sogaard-2025-subword-compositionality] The analysis is purely observational with no causal intervention [peng-chai-sogaard-2025-subword-compositionality]
Structure
Context
subword-compositionality
Confirmed in models
Papers
Understanding Subword Compositionality of Large Language Models — Peng, Qiwei, Chai, Yekun, Søgaard, Anders