MATH · IN · MODELS

A last-token function vector transplants across languages in two stages

measured in 1 paper

Fierro et al. use activation patching and causal-mediation across XGLM-7.5B, EuroLLM-9B, and mT5-XL to disentangle when relation vs language information reaches the last token [fierro-etal-2024-how-do-multilingual-language-models-remember-facts] The last-token representation acts as a function vector that can be causally transplanted into a different-language prompt [fierro-etal-2024-how-do-multilingual-language-models-remember-facts] Transplantation is quantified by the percentage change in the correct object's probability after patching [fierro-etal-2024-how-do-multilingual-language-models-remember-facts] Relation and language information compose in two separable stages rather than being entangled in one representation [fierro-etal-2024-how-do-multilingual-language-models-remember-facts]

Context

a Function Vector causally transplanted across languages rather than across tasks (the setting Function Vectors were originally introduced for), extending the same causal-composition logic to multilingual factual recall, two-stage compositional causal structure (relation information, then language information) validated by patching each stage independently

Papers

How Do Multilingual Language Models Remember Facts? — Fierro, Constanza, Foroutan, Negar, Elliott, Desmond, Søgaard, Anders2024 · arXiv:2410.14387