MATH · IN · MODELS

55 harmfulness-subconcept probes collapse to one dominant direction

measured in 1 paper

Shah et al. train 55 per-subconcept logistic-regression probes (racial hate, weapons, employment scams, etc.) on attention-output states of Llama-3.1-8B-Instruct, replicated on Qwen2-7B-Instruct, each ~0.90 mean accuracy [shah-etal-2025-harmfulness-subconcept-geometry] Stacking the 55 weight vectors, an SVD-based effective rank is K=1 at variance threshold 0.95 for all but the second-to-last layer, so 55 subconcepts share almost one direction [shah-etal-2025-harmfulness-subconcept-geometry] K-means on the weight vectors barely matches the dataset's own category taxonomy (mean Adjusted Rand Index ~3e-4), so the shared direction is not re-deriving the taxonomy [shah-etal-2025-harmfulness-subconcept-geometry] Ablating only the dominant direction matches full-subspace ablation on JailbreakBench safe-rate (~0.91) while raising utility, and steering along it cuts AutoDAN attack success 0.94->0.50 for Llama [shah-etal-2025-harmfulness-subconcept-geometry] Exact effective-rank and jailbreak figures could not be re-extracted from the source (rate-limited), so this entry is deferred for full re-verification, though scope and models are confirmed [shah-etal-2025-harmfulness-subconcept-geometry]

Context

55 per-subconcept logistic-regression probes, SVD-based effective rank (variance-threshold tau), low-rank (near-1D) harmfulness subspace, Adjusted Rand Index vs. dataset taxonomy (near-chance match), subspace vs. dominant-direction ablation, norm-preserving additive steering

Papers

Death by a Thousand Directions: Exploring the Geometry of Harmfulness in LLMs through Subconcept Probing — Shah, McNair, Angeline, Saleena, Kumar, Adhitya Rajendra, Chheda, Naitik, Zhu, Kevin, Sharma, Vasu, O'Brien, Sean, Cai, Will2025 · arXiv:2507.21141