MATH · IN · MODELS

KnowledgeEditor gates the fine-tuning gradient with a hypernetwork outer product

measured in 1 paper

De Cao et al. condition small FFNNs on a bidirectional-LSTM encoding of an edit request to predict per-matrix vectors forming an outer-product gate and bias on the loss gradient [decao-etal-2021-knowledge-editor] A KL-divergence-in-output-space constraint, not a parameter-space norm, prevents collateral damage: an Lp-parameter constraint instead collapses retain accuracy from 98.14 to 45.10 [decao-etal-2021-knowledge-editor] It uses rather than finds geometry, making no claim the edited fact was pre-encoded, and its probe framing concerns only which weight matrices receive large updates [decao-etal-2021-knowledge-editor] Tested on BERT-base (FEVER 98.80% success) and BART-base (zsRE 94.65%), it is the direct architectural predecessor of MEND [decao-etal-2021-knowledge-editor]

Context

hyper-network-predicted outer-product gate and bias on a loss gradient, KL-divergence-in-output-space constraint vs. raw parameter-space norm, USES not FINDS (no claim about pre-existing fact geometry), architectural predecessor of MEND's gradient-factor hypernetwork

Papers

Editing Factual Knowledge in Language Models — De Cao, Nicola, Aziz, Wilker, Titov, Ivan2021 · arXiv:2104.08164