MATH · IN · MODELS

Linguistic category manifolds untangle across transformer layers

measured in 1 paper

Mamou et al. treat each linguistic-label point cloud as an object manifold and use replica mean-field theory (MFTMA) across BERT, RoBERTa, ALBERT, DistilBERT, and GPT-1 [mamou-etal-2020] On masked-token input, manifold capacity (linear separability) consistently increases across layers, driven by joint reduction in manifold radius, dimension, and inter-manifold center correlation [mamou-etal-2020] On unmasked input the effect reverses for word identity while higher-level categories (POS, NER, semantic tags) still gain separability, most strongly for part-of-speech-ambiguous words [mamou-etal-2020]

Context

part-of-speech, manifold capacity, untangling, contextualization, BERTology

Papers

Emergence of Separable Manifolds in Deep Language Representations — Mamou, J., Le, H., Del Rio, M., Stephenson, C., Tang, H., Kim, Y., Chung, S.2020 · arXiv:2006.01095