MATH · IN · MODELS

Knowledge editing shatters transformers' natural cyclic manifold geometry

measured in 1 paper

Nishi et al. confirm via Isomap that pretrained transformers (Llama-3.1-405B-Instruct, GPT-2-Small, Mistral-7B) encode cyclic concepts like months and weekdays as genuine cyclic manifolds, not merely approximate clusters [nishi-etal-2024-representation-shattering-in-transformers] They train a synthetic transformer on a structured knowledge graph and quantify edit-induced distortion with a metric R(D*) = ||D*-D_empty||_F / ||D_empty||_F over pairwise-distance matrices [nishi-etal-2024-representation-shattering-in-transformers] Applying ROME or MEMIT edits measurably shatters the cyclic-manifold structure, distorting the relative positions of non-targeted entities [nishi-etal-2024-representation-shattering-in-transformers] The distortion scales with the edit's counterfactual distance from the targeted fact and degrades factual recall and downstream reasoning [nishi-etal-2024-representation-shattering-in-transformers] The effect is replicated with naturalistic corroboration on real pretrained Llama-3-8B-Instruct and Mamba, not just the synthetic model [nishi-etal-2024-representation-shattering-in-transformers]

Structure

Context

knowledge editing, ROME, MEMIT, cyclic manifolds, collateral distortion, causal validation

Papers

Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing — Nishi, Kento, Ramesh, Rahul, Okawa, Maya, Khona, Mikail, Tanaka, Hidenori, Lubana, Ekdeep Singh2024 · arXiv:2410.17194