MATH · IN · MODELS

Editing MLP output-vector directions suppresses memorized sequences

measured in 1 paper

Hakimi et al. identify MLP neurons implicated in verbatim memorization via logit-lens attribution, then edit each neuron's output vector to add a distractor direction while preserving its other superposed functions [hakimi-etal-2026-output-vector-editing-for-memorization] Four edit modes jointly suppress up to 96.5% of memorized continuations in ensemble, a 2.7x larger effect than zero-ablating the same neurons [hakimi-etal-2026-output-vector-editing-for-memorization] So the direction of the edit, not just the neuron's presence, drives the effect [hakimi-etal-2026-output-vector-editing-for-memorization] About 14% of memorized sequences resist MLP-only editing; ablating top attention heads recovers 60-64% of these, indicating a mechanism split across MLP and attention, and the method transfers to SmolLM-360M, OLMo-1B, and Llama2-7B [hakimi-etal-2026-output-vector-editing-for-memorization]

Context

memorization, directional-editing

Papers

Output Vector Editing for Memorization Mitigation in Large Language Models — Hakimi, Ahmad Dawar, Lei, Kaiwei, Augenstein, Isabelle, Schütze, Hinrich2026 · arXiv:2606.18767