A cross-model universal refusal direction transfers to jailbreak detection
measured in 1 paperYung et al. extend the single-model refusal-direction account to a cross-model claim, defining a universal feature space motivated by model-stitching theory [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction] They identify universal jailbreak features via layer-wise MLP representation propagation and derive a universal refusal direction by averaging per-LLM diff-in-means refusal vectors expressed in this shared space [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction] Because the averaged vector is one-dimensional, transferable jailbreak-prompt detection reduces to a simple linear projection onto it, working in- and out-of-distribution and on black-box models [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction] The universal feature space improves jailbreak detection ~10% over prior single-model baselines; specific model names could not be confirmed from the gated PDF, so the models field is omitted [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction]