55 harmfulness-subconcept probes collapse to one dominant direction
measured in 1 paperShah et al. train 55 per-subconcept logistic-regression probes (racial hate, weapons, employment scams, etc.) on attention-output states of Llama-3.1-8B-Instruct, replicated on Qwen2-7B-Instruct, each ~0.90 mean accuracy [shah-etal-2025-harmfulness-subconcept-geometry] Stacking the 55 weight vectors, an SVD-based effective rank is K=1 at variance threshold 0.95 for all but the second-to-last layer, so 55 subconcepts share almost one direction [shah-etal-2025-harmfulness-subconcept-geometry] K-means on the weight vectors barely matches the dataset's own category taxonomy (mean Adjusted Rand Index ~3e-4), so the shared direction is not re-deriving the taxonomy [shah-etal-2025-harmfulness-subconcept-geometry] Ablating only the dominant direction matches full-subspace ablation on JailbreakBench safe-rate (~0.91) while raising utility, and steering along it cuts AutoDAN attack success 0.94->0.50 for Llama [shah-etal-2025-harmfulness-subconcept-geometry] Exact effective-rank and jailbreak figures could not be re-extracted from the source (rate-limited), so this entry is deferred for full re-verification, though scope and models are confirmed [shah-etal-2025-harmfulness-subconcept-geometry]