MATH · IN · MODELS

Causal-tracing localization does not predict the best edit layer

measured in 1 paper

Hase et al. sweep the target layer for ROME/MEMIT rank-one edits in GPT-J and GPT2-XL and regress edit success on edit layer and causal-tracing effect [hase-etal-2023-does-localization-inform-editing] For ROME on GPT-J, edit layer alone explains 94.7% of edit-success variance while the tracing effect explains 1.6% and adds only 0.1% [hase-etal-2023-does-localization-inform-editing] The raw correlation between edit success and tracing effect is slightly negative (rho=-0.13, p<1e-3), opposite the localize-then-edit assumption [hase-etal-2023-does-localization-inform-editing] A "Fact Forcing" variant reusing tracing's noised-subject input shows a small positive tracing contribution (up to 2.7%), isolated to the shared input format [hase-etal-2023-does-localization-inform-editing] This critiques the layer-selection heuristic, not the rank-one edit mechanism's own geometry [hase-etal-2023-does-localization-inform-editing]

Context

localization (where information is causally represented) vs. editability (where intervention works best) as distinct questions, near-zero/slightly-negative correlation between causal-tracing AIE and rank-one edit success across layers, edit layer choice as by far the dominant predictor of edit success, not localization strength, critique of layer-selection heuristic, not of the outer-product edit mechanism's own geometry

Papers

Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models — Hase, Peter, Bansal, Mohit, Kim, Been, Ghandeharioun, Asma2023 · arXiv:2301.04213