ConceptAttention output-space dot products localize concepts in DiT tokens
measured in 1 paperHelbling et al. thread concept tokens through Flux-Schnell's frozen attention with a one-directional design so concepts read the model's state but never influence the generated image [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features] A saliency map from image-concept dot products in the attention-output space reaches 83.07 accuracy / 71.04 mIoU, versus 74.92/59.90 in cross-attention space and 45.78/29.68 in value space [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features] Concept-image alignment is thus concentrated in the output-projected subspace, beating all 11 compared zero-shot interpretability baselines [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features] Performance concentrates in deeper layers (last 10 of 18) and middle diffusion timesteps, and the method is a passive read-out with no steering claim [helbling-etal-2025-conceptattention-diffusion-transformers-learn-highly-interpretable-features]