MATH · IN · MODELS

Concept-aligned-token SAE feature subgroups in real Gemma-2-2B-IT localize jailbreak susceptibility to mid-to-late layers

measured in 1 paper

Das & Gaur (2026) extract concept-aligned tokens from harmful prompts (BeaverTails, 14 harm categories) in real Gemma-2-2B-IT via cosine similarity to a ReFT-derived harm-concept subspace, then identify Gemma-Scope SAE feature subgroups for those tokens across all 26 layers using three independent grouping strategies (agglomerative clustering, hierarchical-linkage, single-token-driven) [das-gaur-2026-mechanistic-steering-of-llms-reveals-layer-wise-feature-vulnerabilities] All three grouping strategies convergently implicate layers approximately 16-25 as most steerable, and causally amplifying only the top features from an identified subgroup measurably raises an LLM-judged harmfulness score relative to baseline, per category and per layer [das-gaur-2026-mechanistic-steering-of-llms-reveals-layer-wise-feature-vulnerabilities]

Confirmed in models

Method

Papers

Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings — Das, Nilanjana, Gaur, Manas2026 · arXiv:2604.23130