methods / Causal Validation / Causal interventions (steering) / Activation Steering (Addition) / Subspace-Compression Defense (FJLT / Bottleneck Autoencoder)
Subspace-Compression Defense (FJLT / Bottleneck Autoencoder)
Defends against direction-based activation-engineering jailbreaks (ActAdd/Ablation on a linear safety direction) by fine-tuning the model to route its residual stream, or a subset of attention heads, through a low-dimensional linear bottleneck — a Fast Johnson-Lindenstrauss random projection or a learned linear autoencoder — that destroys the exploitable high-dimensional linear separability of the safety concept while preserving enough information for alignment.