MATH · IN · MODELS

Superposed features arrange into uniform polytopes with quantized dimensionality

measured in 1 paper

Elhage et al. train toy ReLU-output autoencoders and find a sharp first-order phase change per feature between not-learned, dedicated-dimension, and superposition [elhage-etal-2022-toy-models-of-superposition] In superposition, features arrange into small uniform polytopes (antipodal pairs, pentagons, tetrahedra, square antiprisms) at quantized sticky fractional-dimensionality values [elhage-etal-2022-toy-models-of-superposition] Adversarial-example vulnerability rises sharply (over 3x) as superposition forms and closely tracks the reciprocal of feature dimensionality [elhage-etal-2022-toy-models-of-superposition] Adversarial training causally reduces superposition, eliminating it entirely only at unreasonably large perturbation budgets (80% input L2 norm) [elhage-etal-2022-toy-models-of-superposition]

Context

superposition (sparse features sharing a dimension), phase change (not-learned / dedicated-dimension / superposition), fractional dimensionality D_i (quantized, "sticky" values), uniform polytopes beyond the simplex (antipodal pair, pentagon, tetrahedron, square antiprism, digon), adversarial-example vulnerability as a causal probe of superposition

Papers

Toy Models of Superposition — Elhage, Nelson, Hume, Tristan, Olsson, Catherine, Schiefer, Nicholas, Henighan, Tom, Kravec, Shauna, Hatfield-Dodds, Zac, Lasenby, Robert, Drain, Dawn, Chen, Carol, Grosse, Roger, McCandlish, Sam, Kaplan, Jared, Amodei, Dario, Wattenberg, Martin, Olah, Christopher2022 · arXiv:2209.10652