MoE routing reflects hidden-state geometry, not domain expertise
measured in 1 paper- Under the load-balancing auxiliary loss, MoE routers are pushed to suppress the shared dominant direction of their hidden states (Proposition 2), so routing tracks hidden-state geometry rather than domain content. [wang-etal-2026-myth-of-expert-specialization] - Hidden-state cosine similarity predicts expert-usage similarity, and two different models solving the same math problem share ~60% of their most-used experts, the same overlap as one model across different questions. [wang-etal-2026-myth-of-expert-specialization] - Prefill expert usage is nearly identical across semantically different inputs and diverges only during generation, with deep-layer router collapse. [wang-etal-2026-myth-of-expert-specialization] - Studied on five production MoE models: gpt-oss-20B, ERNIE-4.5-21B-A3B-Base, Qwen-3-30B, Ling-mini and Trinity-Mini-Base (Moonlight-16B is a comparison only); observational. [wang-etal-2026-myth-of-expert-specialization]