MEMIT batch-edits many facts across MLP layers via least squares
measured in 1 paperMeng et al. show ROME's single rank-one update degrades once applied sequentially for many facts, collapsing by n=10,000 edits [meng-etal-2022-memit] They generalize the linear-associative-memory framing to a batch least-squares objective over many key-value pairs simultaneously, a rank-<=u update [meng-etal-2022-memit] The edit is spread across a range of MLP layers (3-8 in GPT-J), with the residual apportioned equally across remaining layers [meng-etal-2022-memit] At n=10,000 simultaneous edits, MEMIT reaches a CounterFact composite of 85.8 on GPT-J and 82.0 on GPT-NeoX, versus ROME's and MEND's collapse [meng-etal-2022-memit] It uses rather than finds geometry, introducing no new claim about pre-existing key-direction packing [meng-etal-2022-memit]