SD concept directions show composition, style, and texture set at successive denoising steps
measured in 1 paperTinaz et al. train sparse autoencoders on Stable Diffusion v1.4 U-Net activations across the reverse-diffusion trajectory, uncovering human-interpretable concept directions [tinaz-etal-2025-emergence-and-evolution-of-interpretable-concepts-in-diffusion-models] Final scene composition can be predicted from the spatial distribution of activated concepts even before the first denoising step completes [tinaz-etal-2025-emergence-and-evolution-of-interpretable-concepts-in-diffusion-models] Manipulating concept activations confirms a temporal control hierarchy: early denoising steps control composition, mid steps control style once composition is fixed, and late steps affect only minor texture [tinaz-etal-2025-emergence-and-evolution-of-interpretable-concepts-in-diffusion-models]