MATH · IN · MODELS

ConceptAttention output-space dot products localize concepts in DiT tokens

measured in 1 paper

Helbling et al. thread concept tokens through Flux-Schnell's frozen attention with a one-directional design so concepts read the model's state but never influence the generated image [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features] A saliency map from image-concept dot products in the attention-output space reaches 83.07 accuracy / 71.04 mIoU, versus 74.92/59.90 in cross-attention space and 45.78/29.68 in value space [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features] Concept-image alignment is thus concentrated in the output-projected subspace, beating all 11 compared zero-shot interpretability baselines [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features] Performance concentrates in deeper layers (last 10 of 18) and middle diffusion timesteps, and the method is a passive read-out with no steering claim [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features]

Context

one-directional attention operation threading concept tokens through frozen weights, concept residual stream, causally inert with respect to image generation, space ablation (output vs. cross-attention vs. value space) for a dot-product saliency readout, state-of-the-art zero-shot segmentation outperforming 11 baselines, layer-depth and diffusion-timestep concentration of concept-alignment signal

Confirmed in models

Papers

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features — Helbling, Alec, Meral, Tuna Han Salih, Hoover, Ben, Yanardag, Pinar, Chau, Duen Horng2025 · arXiv:2502.04320