Linguistic category manifolds untangle across transformer layers
measured in 1 paperMamou et al. treat each linguistic-label point cloud as an object manifold and use replica mean-field theory (MFTMA) across BERT, RoBERTa, ALBERT, DistilBERT, and GPT-1 [mamou-etal-2020] On masked-token input, manifold capacity (linear separability) consistently increases across layers, driven by joint reduction in manifold radius, dimension, and inter-manifold center correlation [mamou-etal-2020] On unmasked input the effect reverses for word identity while higher-level categories (POS, NER, semantic tags) still gain separability, most strongly for part-of-speech-ambiguous words [mamou-etal-2020]
Structure
Context
part-of-speech, manifold capacity, untangling, contextualization, BERTology
Confirmed in models
Papers
Emergence of Separable Manifolds in Deep Language Representations — Mamou, J., Le, H., Del Rio, M., Stephenson, C., Tang, H., Kim, Y., Chung, S.