Fitted part-of-speech CLIP subspaces disentangle content from appearance
measured in 1 paperOldfield et al. FIT part-of-speech-grounded subspaces in CLIP's joint space using WordNet word lists, via a closed-form class-contrastive trace-maximization objective (a constructed subspace, not one found in free activations) [oldfield-etal-2023-pos-clip-subspaces] Because CLIP embeddings live on a hypersphere, the objective is solved in the tangent space at the intrinsic mean and mapped back, outperforming the flat-Euclidean variant [oldfield-etal-2023-pos-clip-subspaces] Class-invariance metrics confirm each noun/adjective/verb/adverb subspace captures variance nearly exclusively from its own class, beating PCA and Principal Geodesic Analysis [oldfield-etal-2023-pos-clip-subspaces] Projecting a prompt embedding onto the orthogonal complement of the adjective (or artist) subspace before CLIP-conditioned generation selectively removes an artist style or a theme like gore [oldfield-etal-2023-pos-clip-subspaces]