MATH · IN · MODELS

A pose-based sign model handshape geometry matches human perception

measured in 1 paper

Yin et al. compare a pixel-based (I3D) and pose-based (ST-GCN) sign-recognition model on ASL Citizen, testing whether they encode phonological structure rather than shortcuts [yin-etal-2026-sign-language-phonological-perception] On minimal pairs, ST-GCN wins on 81.08% (vs I3D 69.50%), stronger on handshape contrasts while I3D is stronger on location [yin-etal-2026-sign-language-phonological-perception] RSA correlates ST-GCN's handshape similarity with human perceptual confusion (r=0.49) and an articulatory Handshape-Distance metric (r=0.55), versus I3D 0.31/0.20 [yin-etal-2026-sign-language-phonological-perception] The evidence is an RSA distance-correlation to perceptual/articulatory ground truth with no causal intervention; full text was not fully available for deeper verification [yin-etal-2026-sign-language-phonological-perception]

Context

minimal-pairs cosine-similarity discriminability test (intra- vs inter-sign), RSA-style Pearson correlation between model handshape similarity and human perceptual confusion data, RSA-style correlation to an independent articulatory-geometry ground truth (Handshape Distance), architectural dissociation (pose-based models better capture handshape; pixel-based models better capture location)

Papers

Phonological Perception of Sign Language Models — Yin, Kayo, Carter, Jessica, Lu, Alex X., Kocab, Annemarie2026