MATH · IN · MODELS

Vision MoE-CNN experts partition into a stable animate/inanimate split

measured in 1 paper

- A custom contrastive vision MoE-CNN (ResNet backbone, SimCLR objective, ~46M params, E in {4,8,16} experts, top-2 routing) trained from scratch on STL10 develops experts whose representations partition into a stable animate/inanimate split. [tangtartharakul-storrs-2026-vision-moe-representation] - The partition is discovered by RSA: per-expert RDMs over the THINGS images (1,854 concepts) are second-order-correlated and agglomeratively clustered (silhouette-selected), recovering a robust two-cluster animate-vs-inanimate structure. [tangtartharakul-storrs-2026-vision-moe-representation] - The split is stable across all tested expert counts (4, 8, 16) and 10 independently trained seeds, even though top-2 routing does not select the same expert pair, mirroring a known organizational axis of human inferior temporal cortex. [tangtartharakul-storrs-2026-vision-moe-representation] - Nested-CV Lasso against 66 human visual/semantic dimensions shows individual experts are broadly tuned to continuous dimensions (e.g. "animal-related", "movement-related") rather than narrowly domain-specialized; observational, small-scale (STL10, self-supervised only). [tangtartharakul-storrs-2026-vision-moe-representation]

Context

expert-representation RSA clustering, animate/inanimate partition, Lasso-regression tuning-dimension recovery

Papers

Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts — Tangtartharakul, Gene, Storrs, Katherine R.2026 · arXiv:2605.20610