MATH · IN · MODELS

Group-composition networks implement a representation-theoretic Fourier algorithm

measured in 1 paper

Chughtai, Chan & Nanda train small MLPs and transformers on finite-group composition tasks and present a representation-theoretic (Fourier/irrep) algorithm for how such networks implement group composition [chughtai-chan-nanda-2023-toy-model-of-universality] Reverse-engineering real trained weights and logits, plus ablation experiments, confirms networks consistently learn this family of algorithms [chughtai-chan-nanda-2023-toy-model-of-universality] Evidence for universality is mixed: the theory fully characterizes the space of possible circuits, but the specific circuit and its training order vary across differently-initialized networks [chughtai-chan-nanda-2023-toy-model-of-universality] The algorithm covers non-abelian groups (e.g. S5, S6, A5) via higher-dimensional irreps, of which the cyclic circle is only a special sub-case [chughtai-chan-nanda-2023-toy-model-of-universality]

Structure

Context

Fourier/irrep feature geometry, group-composition universality, ablation-confirmed mechanism

Papers

A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations — Chughtai, Bilal, Chan, Lawrence, Nanda, Neel2023 · arXiv:2302.03025