MATH · IN · MODELS

Fitted part-of-speech CLIP subspaces disentangle content from appearance

measured in 1 paper

Oldfield et al. FIT part-of-speech-grounded subspaces in CLIP's joint space using WordNet word lists, via a closed-form class-contrastive trace-maximization objective (a constructed subspace, not one found in free activations) [oldfield-etal-2023-pos-clip-subspaces] Because CLIP embeddings live on a hypersphere, the objective is solved in the tangent space at the intrinsic mean and mapped back, outperforming the flat-Euclidean variant [oldfield-etal-2023-pos-clip-subspaces] Class-invariance metrics confirm each noun/adjective/verb/adverb subspace captures variance nearly exclusively from its own class, beating PCA and Principal Geodesic Analysis [oldfield-etal-2023-pos-clip-subspaces] Projecting a prompt embedding onto the orthogonal complement of the adjective (or artist) subspace before CLIP-conditioned generation selectively removes an artist style or a theme like gore [oldfield-etal-2023-pos-clip-subspaces]

Context

parts of speech, trace maximization, geodesic submanifold, tangent space projection, class invariance, style-content disentanglement, custom visual theme subspace

Papers

Parts of Speech-Grounded Subspaces in Vision-Language Models — Oldfield, James, Tzelepis, Christos, Panagakis, Yannis, Nicolaou, Mihalis A., Patras, Ioannis2023 · arXiv:2305.14053