Dictionary-learned features in Flux.1 causally steer image generation
measured in 1 paperShabalin et al. apply SAEs and Inference-Time Decomposition of Activations to residual-stream embeddings of Flux.1, a large text-to-image diffusion model [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning] SAEs accurately reconstruct Flux.1's residual stream and outperform raw MLP neurons on an automated visual-interpretability pipeline [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning] SAE features causally steer image generation via activation addition, with measured changes in the generated images [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning] ITDA achieves comparable interpretability to SAEs at a different computational tradeoff [shabalin-etal-2025-interpreting-large-text-to-image-diffusion-models-with-dictionary-learning]