MATH · IN · MODELS

MoE router weight vectors geometrically couple to their experts

measured in 1 paper

Ahrac, Hochwald & Geva prove a router weight vector and its matched expert's gate matrix receive gradient updates proportional to the same routed hidden-state direction [ahrac-etal-2026-geometric-coupling-moe] In a real 1B SMoE (9 layers, 64 experts, top-6) trained on ~50B tokens, router score correlates with per-token expert gate-neuron activation at rho=0.43 (p=1.2e-81) [ahrac-etal-2026-geometric-coupling-moe] Auxiliary-loss balancing makes router weight vectors nearly 3x more mutually cosine-similar than loss-free balancing (0.57-0.63 vs 0.13-0.32 across three layers) [ahrac-etal-2026-geometric-coupling-moe] A parameter-free online K-means router achieves the lowest load imbalance while maintaining comparable perplexity [ahrac-etal-2026-geometric-coupling-moe]

Context

router-expert weight coupling, load-balancing loss comparison, centroid router

Papers

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts — Ahrac, Sagi, Hochwald, Noya, Geva, Mor2026 · arXiv:2605.12476