Video-DiT massive activations follow a token-position hierarchy
measured in 1 paper- Video diffusion transformers carry massive-activation channels exceeding 50x the mean activation at a fixed small set of dimensions, consistent across layers and model scales. [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations] - Magnitudes follow a token-position hierarchy: first-latent-frame tokens are largest, with periodic spikes at the VAE's 4x temporal-chunk boundaries, decaying toward the interior over denoising. [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations] - Structured Activation Steering overwrites the top/tail ~8% boundary-token values toward a scaled reference, improving VBench (Wan2.1 81.39 to 81.76, CogVideoX-5B 79.34 to 79.61, Wan2.2 81.74 to 81.93) at negligible cost; a position-agnostic edit instead hurts quality. [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations] - Tested on Wan2.1-T2V-1.3B, Wan2.2-TI2V-5B and CogVideoX-5B (FLUX is an image-DiT baseline only). [cheng-etal-2026-steering-video-diffusion-transformers-massive-activations]