Function-vector heads split into writer and canceller populations with orthogonal OV directions
measured in 1 paperWang shows the function-vector heads identified by Todd et al. in real Pythia (410M-12B), Qwen2.5, and GPT-2-medium are not homogeneous, splitting by sign into writers and cancellers via refined direct logit attribution [wang-2026-function-vector-heads-writers-cancellers] The two sub-populations' mean OV directions are nearly orthogonal (perpendicular fraction 0.96), so cancellers write to a near-orthogonal subspace while still exerting a negative direct causal effect [wang-2026-function-vector-heads-writers-cancellers] Zero-ablating cancellers yields +0.13 to +0.29 nats of logit gain in 6 of 6 main cells, with a consistent +2 to +7 point ICL accuracy effect [wang-2026-function-vector-heads-writers-cancellers] A TOST equivalence test shows cancellers are not simply induction heads in disguise [wang-2026-function-vector-heads-writers-cancellers]