MATH · IN · MODELS

Insecure-code SAE directions sit closer to toxic directions than secure-code ones

measured in 1 paper

Minegishi et al. train SAEs on five models and identify insecure-code, secure-code, and toxic-persona decoder directions [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry] Across all models and layers, the insecure-code direction is consistently more cosine-similar to the toxic-persona direction than the secure-code direction is [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry] So insecure-code finetuning data is geometrically closer to misaligned-persona representations than matched secure-code data [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry] Filtering training data by high insecure-code-direction activation cuts emergent-misalignment behaviors from 87 to 57, beating random removal (84) and an LLM-judge filter (59) [minegishi-etal-2026-understanding-emergent-misalignment-via-feature-superposition-geometry]

Context

cosine similarity between SAE decoder-vector directions as a geometric proxy for shared latent content between a training-data property and a behavioral persona, geometry-based data filtering (removing training examples by their projection onto a feature direction) as a mitigation outperforming both random and LLM-judge baselines

Papers

Understanding Emergent Misalignment via Feature Superposition Geometry — Minegishi, Gouki, Furuta, Hiroki, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka2026 · arXiv:2605.00842