Massive activations provably cause compression valleys synced with attention sinks
measured in 1 paper- Queipo-de-Llano et al. prove (Theorem 1) that massive residual-stream activations necessarily induce representational compression, with tight bounds on entropy reduction and singular-value dominance. [queipodellano-etal-2025-attention-sinks-and-compression-valleys] - Across Pythia 410M and 6.9B training checkpoints (steps 1k-143k), compression valleys, attention sinks and massive activations emerge synchronized around step 1k and persist, at a fixed layer per model. [queipodellano-etal-2025-attention-sinks-and-compression-valleys] - Staged causal ablation of the massive-activation mechanism progressively - and fully, when all stages are ablated - eliminates the compression valley, confirming the causal link; validated further on GPT-OSS-20B and Gemma-7B. [queipodellano-etal-2025-attention-sinks-and-compression-valleys]