MATH · IN · MODELS

Video-DiT massive activations follow a token-position hierarchy

measured in 1 paper

- Video diffusion transformers carry massive-activation channels exceeding 50x the mean activation at a fixed small set of dimensions, consistent across layers and model scales. [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations] - Magnitudes follow a token-position hierarchy: first-latent-frame tokens are largest, with periodic spikes at the VAE's 4x temporal-chunk boundaries, decaying toward the interior over denoising. [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations] - Structured Activation Steering overwrites the top/tail ~8% boundary-token values toward a scaled reference, improving VBench (Wan2.1 81.39 to 81.76, CogVideoX-5B 79.34 to 79.61, Wan2.2 81.74 to 81.93) at negligible cost; a position-agnostic edit instead hurts quality. [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations] - Tested on Wan2.1-T2V-1.3B, Wan2.2-TI2V-5B and CogVideoX-5B (FLUX is an image-DiT baseline only). [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations]

Structure

Context

massive activations as fixed, consistent outlier channel dimensions (rogue-dimension analogue in a new modality), quantified token-position magnitude hierarchy tied to latent temporal chunking, direct overwrite steering of outlier-dimension values as a causal, geometry-tied intervention

Papers

Steering Video Diffusion Transformers with Massive Activations — Cheng, Xianhang, Zheng, Yujian, Xie, Zhenyu, Liao, Tingting, Li, Hao2026 · arXiv:2603.17825