MATH · IN · MODELS

A participation-ratio spectral signal finds attention circuits label-free across scale

measured in 1 paper

Xu identifies attention-head circuits with a three-step recipe: a spectral signal (participation ratio of each head's output activations, integrated over training), a task-pattern screen, and causal ablation with matched-random controls [xu-2026-spectral-probe-circuits] Across 7 configurations (51M-7B, an 8x range, dense and MoE), induction circuits stay 3-6 heads and causally necessary, and the fraction of heads doing specialized computation is conserved at about 17-19% [xu-2026-spectral-probe-circuits] On a 51M TinyStories model, ablating the 4 spectrally-identified circuit heads drops probe accuracy from 0.843 to 0.151 versus no effect for matched-random controls [xu-2026-spectral-probe-circuits] Six pretraining seeds implement the same task with entirely different head sets, and the spectral signal identifies each seed's idiosyncratic circuit without labels, beating nine alternative ranking signals (precision-at-30 0.97 vs 0.93 next-best on Karpathy-124M) [xu-2026-spectral-probe-circuits]

Context

attention-circuits, spectral-signal

Papers

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers — Xu, Yongzhong2026 · arXiv:2605.24059