MATH · IN · MODELS

Linear safety separability emerges above a hidden-dimension threshold

measured in 1 paper

Teo, Abdullaev & Nguyen show via PCA across Qwen0.5B, GPT2-XL, Qwen7B, and Llama2-7B that emotion clusters become distinctly separable only above ~2,000-3,000 hidden dimensions [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment] This emergent linear separability is exactly what white-box jailbreaks exploit, collapsing Llama2-7B-Chat's refusal/safety from 0.97/0.99 to 0.03/0.14 [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment] Two dimension-reduction defenses (FJLT random projection and a learned bottleneck autoencoder) restore refusal/safety to 0.94-0.96, with the bottleneck retaining about 70% of benign helpfulness [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment] A Rademacher-complexity bound explains why shrinking the ambient dimension raises the sample cost of finding a steering direction [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment]

Context

a hidden-dimension threshold (~2,000-3,000) above which an abstract concept becomes linearly separable, the same emergent linear structure being both what makes steering possible and what steering-based jailbreaks exploit, dimension-reduction (random JL projection or learned bottleneck autoencoder) as a defense that destroys exploitable linear separability while preserving task-relevant information

Papers

The Blessing and Curse of Dimensionality in Safety Alignment — Teo, Rachel S.Y., Abdullaev, Laziz U., Nguyen, Tan M.2025 · arXiv:2507.20333