A participation-ratio spectral signal finds attention circuits label-free across scale
measured in 1 paperXu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]