MATH · IN · MODELS

MLLM visual tokens split into sink, dead, alive clusters

measured in 1 paper

- Visual tokens fed into multimodal LLMs partition into three functional clusters, sink (~10%), dead (~30%) and alive (~60%), with near-constant image-agnostic centroids for sink and dead. [fan-etal-2026-visual-tokens-sparsity-redundancy] - Sink-token centroids reach cross-image cosine similarity >0.99 (variance <1e-5) and the largest dead cluster >0.98, a directional collapse of specific functional subgroups. [fan-etal-2026-visual-tokens-sparsity-redundancy] - Pruning all dead tokens raises accuracy +1.0 while removing the same number of random tokens loses -2.6, confirming the typing causally (Table 2, LLaVA-1.5-7B). [fan-etal-2026-visual-tokens-sparsity-redundancy] - Primary model LLaVA-1.5-7B, with generalization to LLaVA-1.5-13B, InternVL-3-8B and Qwen2.5-VL. [fan-etal-2026-visual-tokens-sparsity-redundancy]

Structure

Context

sink/dead/alive visual-token clusters, cross-image centroid cosine similarity, cluster-targeted ablation

Papers

What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models — Fan, Yingqi, Tong, Junlong, Zhao, Anhao, Shen, Xiaoyu2026 · arXiv:2603.00510