MATH · IN · MODELS

Custom research sparse mixture-of-experts transformer

Structures found in this family (1)

By model (2)

Custom 1B-parameter sparse MoE transformer (Ahrac et al. 2026) · 1B, 9 layers, 64 experts, top-6 routing, 2 shared experts
Semantic Trajectory MoE, 'Marathon' configuration (Ternovtsii & Bilak 2026) · 76-84M, 8 layers, 1024 rank-1 experts/layer, cosine-routing (d_space=64), top-K=4

Papers

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts (2026), Geometric Routing Enables Causal Expert Control in Mixture of Experts (2026)