MATH · IN · MODELS

Function-vector heads split into writer and canceller populations with orthogonal OV directions

measured in 1 paper

Wang shows the function-vector heads identified by Todd et al. in real Pythia (410M-12B), Qwen2.5, and GPT-2-medium are not homogeneous, splitting by sign into writers and cancellers via refined direct logit attribution [wang-2026-function-vector-heads-writers-cancellers] The two sub-populations' mean OV directions are nearly orthogonal (perpendicular fraction 0.96), so cancellers write to a near-orthogonal subspace while still exerting a negative direct causal effect [wang-2026-function-vector-heads-writers-cancellers] Zero-ablating cancellers yields +0.13 to +0.29 nats of logit gain in 6 of 6 main cells, with a consistent +2 to +7 point ICL accuracy effect [wang-2026-function-vector-heads-writers-cancellers] A TOST equivalence test shows cancellers are not simply induction heads in disguise [wang-2026-function-vector-heads-writers-cancellers]

Context

writer/canceller FV-head sub-populations, near-orthogonal OV directions, refined direct logit attribution

Papers

Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning — Wang, Han-yu2026 · arXiv:2606.07560