Custom research sparse mixture-of-experts transformer
Structures found in this family (1)
By model (2)
Custom 1B-parameter sparse MoE transformer (Ahrac et al. 2026) · 1B, 9 layers, 64 experts, top-6 routing, 2 shared experts
Semantic Trajectory MoE, 'Marathon' configuration (Ternovtsii & Bilak 2026) · 76-84M, 8 layers, 1024 rank-1 experts/layer, cosine-routing (d_space=64), top-K=4
Observations (2)
Papers
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts (2026), Geometric Routing Enables Causal Expert Control in Mixture of Experts (2026)