MATH · IN · MODELS

A cross-model universal refusal direction transfers to jailbreak detection

measured in 1 paper

Yung et al. extend the single-model refusal-direction account to a cross-model claim, defining a universal feature space motivated by model-stitching theory [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction] They identify universal jailbreak features via layer-wise MLP representation propagation and derive a universal refusal direction by averaging per-LLM diff-in-means refusal vectors expressed in this shared space [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction] Because the averaged vector is one-dimensional, transferable jailbreak-prompt detection reduces to a simple linear projection onto it, working in- and out-of-distribution and on black-box models [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction] The universal feature space improves jailbreak detection ~10% over prior single-model baselines; specific model names could not be confirmed from the gated PDF, so the models field is omitted [yung-etal-2026-characterising-universal-jailbreak-features-refusal-direction]

Context

universal jailbreak feature space motivated by model-stitching theory, universal refusal direction as the average of per-LLM diff-in-means refusal vectors, one-dimensional linear projection for transferable jailbreak detection, transfer to out-of-distribution and black-box models

Confirmed in models

Papers

Characterising Universal Jailbreak Features and Refusal Direction in LLMs — Yung, Canaan, Huang, Hanxun, Leckie, Christopher, Erfani, Sarah Monazam2026