Insecure-code SAE directions sit closer to toxic directions than secure-code ones
measured in 1 paperMinegishi et al. train SAEs on five models and identify insecure-code, secure-code, and toxic-persona decoder directions [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry] Across all models and layers, the insecure-code direction is consistently more cosine-similar to the toxic-persona direction than the secure-code direction is [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry] So insecure-code finetuning data is geometrically closer to misaligned-persona representations than matched secure-code data [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry] Filtering training data by high insecure-code-direction activation cuts emergent-misalignment behaviors from 87 to 57, beating random removal (84) and an LLM-judge filter (59) [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry]