A last-token function vector transplants across languages in two stages
measured in 1 paperFierro et al. use activation patching and causal-mediation across XGLM-7.5B, EuroLLM-9B, and mT5-XL to disentangle when relation vs language information reaches the last token [fierro-etal-2024-how-do-multilingual-language-models-remember-facts] The last-token representation acts as a function vector that can be causally transplanted into a different-language prompt [fierro-etal-2024-how-do-multilingual-language-models-remember-facts] Transplantation is quantified by the percentage change in the correct object's probability after patching [fierro-etal-2024-how-do-multilingual-language-models-remember-facts] Relation and language information compose in two separable stages rather than being entangled in one representation [fierro-etal-2024-how-do-multilingual-language-models-remember-facts]