MATH · IN · MODELS

Verb aspect is encoded as linear semantic-scale directions

measured in 1 paper

Li, Chersoni & Hsu apply Grand et al.'s semantic-projection method to build stativity, telicity, and durativity scale directions from corpus-averaged verb embeddings in BERT-base-uncased and GPT-2-small [li-chersoni-hsu-2024] Verb activations projected onto each scale show significant extreme-group differences (Mann-Whitney U, p<0.05) across layers [li-chersoni-hsu-2024] Stativity is most robustly encoded, telicity is significant in all BERT layers, and durativity is least consistent [li-chersoni-hsu-2024] Later layers of both models show declining separation, attributed to rising anisotropy that is stronger in GPT-2 [li-chersoni-hsu-2024] A cosine-similarity test of the Imperfective Paradox holds only in BERT's early layers, while GPT-2 shows the reverse pattern under high anisotropy [li-chersoni-hsu-2024]

Context

semantic projection, Vendler verb classes, Imperfective Paradox, anisotropy interaction

Papers

Investigating Aspect Features in Contextualized Embeddings with Semantic Scales and Distributional Similarity — Li, Yuxi, Chersoni, Emmanuele, Hsu, Yu-Yin2024