Beta-VAE axis-alignment emerges from information-bottleneck pressure
measured in 1 paperBurgess et al. give a rate-distortion account of why beta-VAE's latent code becomes axis-aligned with a dataset's independent generative factors [burgess-etal-2018-understanding-disentangling] The beta-weighted KL term upper-bounds per-channel capacity, forcing data locality, while diagonal-covariance allocation drives factors onto separate axes [burgess-etal-2018-understanding-disentangling] As target capacity C rises from 0.5 to 25 nats, per-factor KL is allocated in a fixed order (position, then scale, shape, rotation) [burgess-etal-2018-understanding-disentangling] Latent traversals causally isolate each top-KL dimension's effect to exactly one factor, while lowest-KL dimensions have no effect [burgess-etal-2018-understanding-disentangling] A standard VAE (beta=1) shows the same local smoothness but fragmented, non-axis-aligned factor coding [burgess-etal-2018-understanding-disentangling]