MATH · IN · MODELS

Dictionary-learned features in Flux.1 causally steer image generation

measured in 1 paper

Shabalin et al. apply SAEs and Inference-Time Decomposition of Activations to residual-stream embeddings of Flux.1, a large text-to-image diffusion model [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning] SAEs accurately reconstruct Flux.1's residual stream and outperform raw MLP neurons on an automated visual-interpretability pipeline [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning] SAE features causally steer image generation via activation addition, with measured changes in the generated images [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning] ITDA achieves comparable interpretability to SAEs at a different computational tradeoff [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning]

Context

diffusion-models, dictionary-learning

Confirmed in models

Papers

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning — Shabalin, Stepan, Panda, Ayush, Kharlapenko, Dmitrii, Ali, Abdur Raheem, Hao, Yixiong, Conmy, Arthur2025 · arXiv:2505.24360