MATH · IN · MODELS

ROME localizes facts to a two-peak causal site and edits via a rank-one direction

measured in 1 paper

Meng et al. use causal tracing across 1,000 factual statements on GPT-2 XL to map a strongly causal early MLP site at the subject's last token, distinct from a late attention site at the final token [meng-etal-2022-rome] A severed-module variant confirms the early site depends specifically on MLP computation [meng-etal-2022-rome] They model the responsible MLP's down-projection as a linear associative memory and derive a closed-form rank-one update inserting one key-value association as an outer-product direction [meng-etal-2022-rome] It uses rather than finds geometry (the value is optimized per edit), reaching CounterFact composites of 89.2 (GPT-2 XL) and 91.5 (GPT-J) with high neighborhood specificity [meng-etal-2022-rome]

Context

two-peak (early MLP / late attention) causal localization of facts, linear associative memory (WK≈V) view of MLP weight matrices, rank-one closed-form weight update as an outer-product direction mechanism, USES not FINDS (edited value is optimized, not extracted from existing geometry)

Papers

Locating and Editing Factual Associations in GPT — Meng, Kevin, Bau, David, Andonian, Alex, Belinkov, Yonatan2022 · arXiv:2202.05262