MATH · IN · MODELS

PhonSSM's imposed orthogonal phonological subspaces are causally load-bearing

measured in 1 paper

Cheng, Jin & Zhang build PhonSSM, whose Phonological Decomposition Module projects sign features via four parallel MLPs into four 32-dim component subspaces (handshape/location/movement/orientation) [cheng-jin-zhang-2026-phonssm] The four subspaces are made orthogonal by an explicit orthogonality loss (imposed, not discovered): mean pairwise cosine similarity is 0.12 with the loss versus 0.67 without it [cheng-jin-zhang-2026-phonssm] Linear probes show clean dissociation (handshape branch 78.4% handshape vs 31.2% location, chance 8.3%) [cheng-jin-zhang-2026-phonssm] For 47 minimal pairs, swapping only the differing component embedding flips the prediction to the pair partner 73.2% of the time versus 12.4% for control swaps (p<0.001) [cheng-jin-zhang-2026-phonssm] PhonSSM reaches 72.1% on WLASL2000 (+18.4pp over skeleton SOTA) and 53.34% on a 5,565-sign merged dataset [cheng-jin-zhang-2026-phonssm]

Context

four orthogonal phonological-component subspaces (handshape/location/movement/orientation), orthogonality loss and measured cosine-similarity separation (0.12 vs. 0.67), linear-probe dissociation across component subspaces, minimal-pair component-swap causal intervention (73.2% vs. 12.4% flip rate, p<0.001), vocabulary-scale (5,565-sign) skeleton-based recognition benchmark

Papers

State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition — Cheng, Bryan, Jin, Austin, Zhang, Jasper2026 · arXiv:2604.08761