A pose-based sign model handshape geometry matches human perception
measured in 1 paperYin et al. compare a pixel-based (I3D) and pose-based (ST-GCN) sign-recognition model on ASL Citizen, testing whether they encode phonological structure rather than shortcuts [yin-etal-2026-sign-language-phonological-perception] On minimal pairs, ST-GCN wins on 81.08% (vs I3D 69.50%), stronger on handshape contrasts while I3D is stronger on location [yin-etal-2026-sign-language-phonological-perception] RSA correlates ST-GCN's handshape similarity with human perceptual confusion (r=0.49) and an articulatory Handshape-Distance metric (r=0.55), versus I3D 0.31/0.20 [yin-etal-2026-sign-language-phonological-perception] The evidence is an RSA distance-correlation to perceptual/articulatory ground truth with no causal intervention; full text was not fully available for deeper verification [yin-etal-2026-sign-language-phonological-perception]