MATH · IN · MODELS

Multi-feature SAE steering vectors force instruction-following

measured in 1 paper

He et al. identify instruction-relevant SAE latents via a sensitivity score, then build multi-feature steering vectors from Gemma Scope / Llama Scope decoder directions [he-etal-2025-saif-a-sparse-autoencoder-framework-for-interpreting-and-steering-instruction-following-of-language-models] Adding them to the residual stream forces instruction-following (translation/summarization/keyword) at over 30% strict and up to ~0.7 loose accuracy versus near-zero for single-latent steering, optimal at k=15 latents [he-etal-2025-saif-a-sparse-autoencoder-framework-for-interpreting-and-steering-instruction-following-of-language-models] Last-layer placement is critical (Gemma-2-2b-it loose accuracy 0.64 at layer 25 drops to 0.33 by layer 24), and post-instruction positioning beats pre-instruction [he-etal-2025-saif-a-sparse-autoencoder-framework-for-interpreting-and-steering-instruction-following-of-language-models]

Context

instruction-following, sparse-autoencoders

Papers

SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models — He, Yifei, Zhao, Xiaotian, Qiao, Yuxuan, Yang, Yi, Payani, Ali, Ma, Zhichao, Du, Mengnan2025 · arXiv:2502.11356