MATH · IN · MODELS

Speech models encode neighbor-phone info in orthogonal positional subspaces

measured in 1 paper

Choi et al. show a single S3M frame (wav2vec 2.0, HuBERT, WavLM Large) encodes phonological vectors not only for the current phone but for its neighbors, extracted via difference-of-means at relative positions -2 to +2 [choi-etal-2026-position-dependent-orthogonal-subspaces] Cosine similarity between vectors from different relative positions is substantially lower than within the same position across 8 phonological features, an orthogonality-of-subspaces structure preserved layer-by-layer [choi-etal-2026-position-dependent-orthogonal-subspaces] Vector norm decreases monotonically with distance from the center phone, with WavLM showing the clearest trapezoidal effective context window [choi-etal-2026-position-dependent-orthogonal-subspaces] The position-dependent subspace in use switches at annotated TIMIT phonetic boundaries rather than a fixed temporal window; no causal intervention is performed [choi-etal-2026-position-dependent-orthogonal-subspaces]

Context

phonological vector arithmetic, center-pooled frame representation, positional orthogonality, effective context window, phonetic boundary alignment

Papers

Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces — Choi, Kwanghee, Yeo, Eunjung, Cho, Cheol Jun, Mortensen, David R., Harwath, David2026 · arXiv:2603.12642