MATH · IN · MODELS

A backdoor trigger hides orthogonal to the probed language direction

measured in 1 paper

Kulumba et al. plant a 9-token Latin trigger during pretraining of Gaperon-8B that switches generation from English to French, then decompose a 3-phase causal circuit via activation patching [kulumba-etal-2026-language-switching-triggers-take-a-latent-detour-through-language-models] In the critical phase, linear language-identity probes classify the representation as English throughout mid-to-late layers even though patching confirms the trigger is causally present [kulumba-etal-2026-language-switching-triggers-take-a-latent-detour-through-language-models] The trigger occupies a subspace orthogonal to the natural language-identity direction, converted to French only by a final-layer MLP that accounts for approximately 62% (plus or minus 8%) of the total causal effect [kulumba-etal-2026-language-switching-triggers-take-a-latent-detour-through-language-models]

Context

a causally-present signal that a linear probe fails to detect because it lives in a subspace orthogonal to the probed direction, distinguishing "probeable" from "causally present" more sharply than in most probing papers, a multi-phase activation-patching decomposition localizing where an orthogonal-subspace signal gets converted into the probed direction

Confirmed in models

Papers

Language-Switching Triggers Take a Latent Detour Through Language Models — Kulumba, Frank, Antoun, Wissam, Lasnier, Elyas, Sagot, Benoît, Seddah, Djamé2026 · arXiv:2605.18646