MATH · IN · MODELS

A fitted linear projection (PISCO) disentangles style from content

measured in 1 paper

Ngweta et al. show pretrained ResNet-18 (supervised, ImageNet) and SimCLR ResNet-18 (CIFAR-10) features are a linearly-entangled mixture of style and content factors [ngweta-etal-2023-simple-disentanglement-of-style-and-content-in-visual-representations] They derive PISCO, a provably-correct linear projection fitted post-hoc (via two theorems) to separate style and content with no retraining, a constructed disentangling map rather than a structure discovered in free activations [ngweta-etal-2023-simple-disentanglement-of-style-and-content-in-visual-representations] On CIFAR-10 PISCO raises style-correlation recovery (e.g. rotation: SimCLR 0.368->0.945) and cuts a disentanglement-discrepancy metric (rotation 0.212->0.029) [ngweta-etal-2023-simple-disentanglement-of-style-and-content-in-visual-representations] Discarding the isolated style-subspace factors improves out-of-distribution accuracy under spurious style shift while preserving in-distribution accuracy [ngweta-etal-2023-simple-disentanglement-of-style-and-content-in-visual-representations]

Context

disentanglement, style-content-separation

Papers

Simple Disentanglement of Style and Content in Visual Representations — Ngweta, Lilian, Maity, Subha, Gittens, Alex, Sun, Yuekai, Yurochkin, Mikhail2023 · arXiv:2302.09795