MATH · IN · MODELS

Massive activations provably cause compression valleys synced with attention sinks

measured in 1 paper

- Queipo-de-Llano et al. prove (Theorem 1) that massive residual-stream activations necessarily induce representational compression, with tight bounds on entropy reduction and singular-value dominance. [queipodellano-etal-2025-attention-sinks-and-compression-valleys] - Across Pythia 410M and 6.9B training checkpoints (steps 1k-143k), compression valleys, attention sinks and massive activations emerge synchronized around step 1k and persist, at a fixed layer per model. [queipodellano-etal-2025-attention-sinks-and-compression-valleys] - Staged causal ablation of the massive-activation mechanism progressively - and fully, when all stages are ablated - eliminates the compression valley, confirming the causal link; validated further on GPT-OSS-20B and Gemma-7B. [queipodellano-etal-2025-attention-sinks-and-compression-valleys]

Context

attention-sinks, massive-activations, representational-compression

Papers

Attention Sinks and Compression Valleys in LLMs Are Two Sides of the Same Coin — Queipo-de-Llano, Enrique, Arroyo, Álvaro, Barbero, Federico, Dong, Xiaowen, Bronstein, Michael, LeCun, Yann, Shwartz-Ziv, Ravid2025 · arXiv:2510.06477