MATH · IN · MODELS

Diffusion h-space supports linear, timestep-consistent edit directions

measured in 1 paper

Kwon, Jeong & Uh identify h-space, the deepest U-Net bottleneck of a frozen diffusion model, as a linear semantic latent space [kwon-etal-2022-asyrp-h-space] Edit directions are found by optimizing a per-timestep perturbation via a directional CLIP loss, with their asymmetric reverse process (Asyrp) confining editing to h_t [kwon-etal-2022-asyrp-h-space] The directions are homogeneous (one image's direction transfers to others), linear (scaling and additive composition work), and robust to matched-magnitude h-space noise [kwon-etal-2022-asyrp-h-space] A single time-invariant global direction closely reproduces the per-timestep editing effect [kwon-etal-2022-asyrp-h-space] An 80-participant study prefers Asyrp over DiffusionCLIP for quality in 98.36% of trials, across three architectures and five datasets with frozen checkpoints [kwon-etal-2022-asyrp-h-space]

Context

h-space, diffusion bottleneck, semantic image editing, directional CLIP loss, asymmetric reverse process, editing strength, quality deficiency, homogeneity, linearity, robustness across timesteps

Papers

Diffusion Models already have a Semantic Latent Space — Kwon, Mingi, Jeong, Jaeseok, Uh, Youngjung2022 · arXiv:2210.10960