Bottleneck-localized steering beats full-network steering in music diffusion
measured in 1 paperStaniszewski et al. use activation patching to localize a semantic bottleneck across three text-to-music diffusion models (AudioLDM2, Ace-Step, Stable Audio Open) [staniszewski-etal-2026-tada-tuning-audio-diffusion-models-through-activation-steering] Steering only at these 2-4 bottleneck layers, via mean-difference direction vectors or TopK SAE decoder columns, beats full-network steering [staniszewski-etal-2026-tada-tuning-audio-diffusion-models-through-activation-steering] On Ace-Step "Piano," LPAPS drops to 2.188 for layers {6,7} versus 2.707 for all 24 layers [staniszewski-etal-2026-tada-tuning-audio-diffusion-models-through-activation-steering] SAE-based bottleneck steering achieves the best preservation (LPAPS 1.973) [staniszewski-etal-2026-tada-tuning-audio-diffusion-models-through-activation-steering]