MATH · IN · MODELS

CCA and mutual-information trajectories differ by training objective across depth

measured in 1 paper

Voita et al. track layer-wise CCA similarity (against each model's own final layer) and mutual information with past/future tokens and surface identity across three training objectives (MT, LM, MLM) [voita-etal-2019-bottom-up-evolution-of-representations] LM representations show monotonically increasing divergence from early layers, discarding past-token information while building toward the future-prediction target [voita-etal-2019-bottom-up-evolution-of-representations] MLM representations first move away from surface token identity in early-to-middle layers, then partially reconstruct it near the output [voita-etal-2019-bottom-up-evolution-of-representations] The trajectory shape is thus systematically training-objective-dependent; no causal intervention is performed [voita-etal-2019-bottom-up-evolution-of-representations]

Context

CCA similarity trajectory, mutual information across depth, training-objective-dependent geometry

Papers

The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives — Voita, Elena, Sennrich, Rico, Titov, Ivan2019 · arXiv:1909.01380