Speech SSM embeddings link words to inflected forms linearly
measured in 1 paper- Wav2Vec2 speech representations exhibit a global linear-offset geometry: a single mean difference vector links English base nouns/verbs to their regular inflected forms, so b-a+c recovers the correct inflected target near the top rank. [gauthier-etal-2025-emergent-morphophonology] - Against a random-chance rank of 94,306, a word-optimized probe reaches mean rank ~1 (noun plurals) and ~7.9 (verb 3SG) at layer 8. [gauthier-etal-2025-emergent-morphophonology] - The geometry reflects lexical distributional regularities rather than explicit phonological allomorphs; the probe strips morphological/phonological sensitivity while keeping the linear offset (morphology-mismatch rank difference 7.1 vs 34.7 for the raw model). [gauthier-etal-2025-emergent-morphophonology] - The model is facebook/wav2vec2-base (pretrained on 960h LibriSpeech); English noun-plural and 3SG-verb inflection only; observational. [gauthier-etal-2025-emergent-morphophonology]