Parts of Speech-Grounded Subspaces in Vision-Language Models
James Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou, Ioannis Patras
摘要
Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased towards specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP's joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What's more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists' painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists' styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Instance-Level Composed Image RetrievalBill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis 等NeurIPS 2025 · 被引用 14 次
- On the rankability of visual embeddingsAnkit Sonthalia, Arnas Uselis, Seong Joon OhNeurIPS 2025 · 被引用 4 次
- Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsDavide Berasi, Matteo Farina, Massimiliano Mancini, Elisa Ricci 等CVPR 2025
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Disentangling visual and written concepts in CLIPJoanna Materzynska, Antonio Torralba, David BauCVPR 2022 · 被引用 32 次
- Attribute-based Visual Reprogramming for Vision-Language ModelsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi 等ICLR 2025
- CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable and Controllable Text-Guided Face ManipulationChenliang Zhou, Fangcheng Zhong, Cengiz ÖztireliSIGGRAPH 2023 · 被引用 13 次
- Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language ModelZipeng Xu, Tianwei Lin, Hao Tang, Fu Li 等CVPR 2022 · 被引用 38 次
- Language-biased image classification: evaluation based on semantic representationsYoann Lemesle, Masataka Sawayama, Guillermo Valle Pérez, Maxime Adolphe 等ICLR 2022 · 被引用 8 次
