MATH · IN · MODELS
methods / Causal Validation / Causal interventions (steering) / Activation Steering (Addition) / Subspace-Compression Defense (FJLT / Bottleneck Autoencoder)

Subspace-Compression Defense (FJLT / Bottleneck Autoencoder)

Techniqueadvanced

Defends against direction-based activation-engineering jailbreaks (ActAdd/Ablation on a linear safety direction) by fine-tuning the model to route its residual stream, or a subset of attention heads, through a low-dimensional linear bottleneck — a Fast Johnson-Lindenstrauss random projection or a learned linear autoencoder — that destroys the exploitable high-dimensional linear separability of the safety concept while preserving enough information for alignment.

Used in (1 observation)

structure: Linear Separability · models: Qwen 0.5B (base), GPT-2 XL, Qwen 7B (base), Llama-2-7B, Llama-2-7B-Chat, Gemma 1.1 7B Instruct, Qwen2-7B-Instruct · paper: The Blessing and Curse of Dimensionality in Safety Alignment