Superposed features arrange into uniform polytopes with quantized dimensionality
measured in 1 paperElhage et al. train toy ReLU-output autoencoders and find a sharp first-order phase change per feature between not-learned, dedicated-dimension, and superposition [elhage-etal-2022-toy-models-of-superposition] In superposition, features arrange into small uniform polytopes (antipodal pairs, pentagons, tetrahedra, square antiprisms) at quantized sticky fractional-dimensionality values [elhage-etal-2022-toy-models-of-superposition] Adversarial-example vulnerability rises sharply (over 3x) as superposition forms and closely tracks the reciprocal of feature dimensionality [elhage-etal-2022-toy-models-of-superposition] Adversarial training causally reduces superposition, eliminating it entirely only at unreasonably large perturbation budgets (80% input L2 norm) [elhage-etal-2022-toy-models-of-superposition]