Linear safety separability emerges above a hidden-dimension threshold
measured in 1 paperTeo, Abdullaev & Nguyen show via PCA across Qwen0.5B, GPT2-XL, Qwen7B, and Llama2-7B that emotion clusters become distinctly separable only above ~2,000-3,000 hidden dimensions [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment] This emergent linear separability is exactly what white-box jailbreaks exploit, collapsing Llama2-7B-Chat's refusal/safety from 0.97/0.99 to 0.03/0.14 [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment] Two dimension-reduction defenses (FJLT random projection and a learned bottleneck autoencoder) restore refusal/safety to 0.94-0.96, with the bottleneck retaining about 70% of benign helpfulness [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment] A Rademacher-complexity bound explains why shrinking the ambient dimension raises the sample cost of finding a steering direction [teo-etal-2025-blessing-and-curse-of-dimensionality-in-safety-alignment]