MATH · IN · MODELS

MoE experts are functionally decorrelated but only partly subspace-separated

measured in 1 paper

Liu measures two pretrained sparse MoE models, Mixtral-8x7B (8 experts/layer, top-2) and Qwen1.5-MoE-A2.7B (60 experts/layer, top-4) [liu-2026-geometric-asymmetry-moe-specialization] Cross-expert Jacobian cosine similarity clusters near zero (Mistral middle-layer mean 0.062; Qwen ~0.000-0.001), showing experts are strongly functionally decorrelated [liu-2026-geometric-asymmetry-moe-specialization] Yet their top-5-PCA subspaces sit at Grassmannian distances (2.06-2.69) well below the theoretical maximum ~3.51, so subspaces are distinct but only partially separated [liu-2026-geometric-asymmetry-moe-specialization] A from-scratch 8-expert Transformer isolates routing's causal role: mean Grassmann distance is 2.463 under top-k routing vs 0.480 under fully-soft routing [liu-2026-geometric-asymmetry-moe-specialization]

Context

cross-expert Jacobian decorrelation, Grassmannian subspace distance between experts, routing-sparsity ablation

Papers

Geometric Asymmetry in MoE Specialization: Functional Decorrelation and Representational Overlap — Liu, Feilong2026 · arXiv:2605.16349