MATH · IN · MODELS

Massive activations re-emerge under a protected residual stream and resist a stronger sparsity penalty

measured in 1 paper

Vemula trains real decoder-only transformers from scratch on FineWeb-Edu at 160M and 290M parameters under four matched configurations (vanilla, QK-normalized "suppress", a novel Ledger-Residuals architecture splitting the residual stream into a mutable Deliberation stream and a protected Commitment stream, and a Commitment variant with 6x stronger commit-sparsity penalty), measuring massive activations via excess kurtosis, fixed-dimension ratio, dominant-dimension persistence, and start-token concentration [vemula-2026-massive-activations-architecturally-robust] At 290M parameters and matched validation loss, the vanilla transformer shows a massive-activation channel at dimension 828 (persistence 0.92, fixed-dimension ratio 28.3, start-token concentration 1.01), which QK-normalization essentially removes (ratio 6.8, concentration 0.13), but the Ledger-Residuals commitment channel rebuilds a massive activation at a different dimension (dimension 22, persistence 0.58, ratio 7.0, concentration 3.36) despite architecturally isolating it from the mutable stream [vemula-2026-massive-activations-architecturally-robust] A 6x stronger commit-sparsity penalty on the same Ledger-Residuals architecture makes the re-emerged massive activation worse rather than better (persistence rises to 0.96, concentration to 4.69), and the same qualitative re-emergence pattern replicates at 160M parameters (vanilla kurtosis 81, ratio 34, persistence 54%; ledger commitment channel rebuilds at ratio 13, concentration 3.04) [vemula-2026-massive-activations-architecturally-robust]

Method

Papers

Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test — Vemula, Maruthi2026 · arXiv:2606.20743