CCA and mutual-information trajectories differ by training objective across depth
measured in 1 paperVoita et al. track layer-wise CCA similarity (against each model's own final layer) and mutual information with past/future tokens and surface identity across three training objectives (MT, LM, MLM) [voita-etal-2019-bottom-up-evolution-of-representations] LM representations show monotonically increasing divergence from early layers, discarding past-token information while building toward the future-prediction target [voita-etal-2019-bottom-up-evolution-of-representations] MLM representations first move away from surface token identity in early-to-middle layers, then partially reconstruct it near the output [voita-etal-2019-bottom-up-evolution-of-representations] The trajectory shape is thus systematically training-objective-dependent; no causal intervention is performed [voita-etal-2019-bottom-up-evolution-of-representations]