A backdoor trigger hides orthogonal to the probed language direction
measured in 1 paperKulumba et al. plant a 9-token Latin trigger during pretraining of Gaperon-8B that switches generation from English to French, then decompose a 3-phase causal circuit via activation patching [kulumba-etal-2026-language-switching-triggers-take-a-latent-detour-through-language-models] In the critical phase, linear language-identity probes classify the representation as English throughout mid-to-late layers even though patching confirms the trigger is causally present [kulumba-etal-2026-language-switching-triggers-take-a-latent-detour-through-language-models] The trigger occupies a subspace orthogonal to the natural language-identity direction, converted to French only by a final-layer MLP that accounts for approximately 62% (plus or minus 8%) of the total causal effect [kulumba-etal-2026-language-switching-triggers-take-a-latent-detour-through-language-models]