MATH · IN · MODELS

SwiGLU spikes collapse keys into a sink-forming subspace

measured in 1 paper

- SwiGLU feed-forward blocks act as directional quadratic amplifiers, creating massive activations for tokens aligned with a shared "spike direction". [sun-etal-2026-the-spike-the-sparse-and-the-sink] - Pre-norm RMSNorm's bounded-range property (Theorem B.3) collapses these spike tokens' normalized key vectors into a low-dimensional near-constant subspace, and sink heads' query subspaces align with it - explaining attention sinks as geometric alignment, not semantic relevance. [sun-etal-2026-the-spike-the-sparse-and-the-sink] - Causal ablations on a controlled 7B model trained from scratch show spikes and sinks are separable: DynamicTanh eliminates spikes without hurting sinks or perplexity, and long-context-only training sharply reduces the sink ratio. [sun-etal-2026-the-spike-the-sparse-and-the-sink] - Broader analysis applied to pre-norm Llama-family models (approximated here by Llama-2-7B). [sun-etal-2026-the-spike-the-sparse-and-the-sink]

Context

attention-sinks, massive-activations, swiglu

Papers

The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks — Sun, Shangwen, Canziani, Alfredo, LeCun, Yann, Zhu, Jiachen2026 · arXiv:2603.05498