Disentangling visual and written concepts in CLIP
Joanna Materzynska, Antonio Torralba, David Bau
Abstract
The CLIP network measures the similarity between natural text and images; in this work, we investigate the entanglement of the representation of word images and natural images in its image encoder. First, we find that the image encoder has an ability to match word images with natural images of scenes described by those words. This is consistent with previous research that suggests that the meaning and the spelling of a word might be entangled deep within the network. On the other hand, we also find that CLIP has a strong ability to match nonsense words, suggesting that processing of letters is separated from processing of their meaning. To explicitly determine whether the spelling capability of CLIP is separable, we devise a procedure for identifying representation subspaces that selectively isolate or eliminate spelling capabilities. We benchmark our methods against a range of retrieval tasks, and we also test them by measuring the appearance of text in CLIP-guided generated images. We find that our methods are able to cleanly separate spelling capabilities of CLIP from the visual processing of natural images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99e75d37-e311-4046-a4c1-a9784a914ef4Cited by top-tier papers32
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- Interpreting CLIP's Image Representation via Text-Based DecompositionYossi Gandelsman, Alexei A. Efros, Jacob SteinhardtICLR 2024 · 179 citations
- What the DAAM: Interpreting Stable Diffusion Using Cross AttentionRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang et al.ACL 2023 · 93 citations
- Decomposing and Editing Predictions by Modeling Model ComputationHarshay Shah, Andrew Ilyas, Aleksander MadryICML 2024 · 25 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
Related papers
- Parts of Speech-Grounded Subspaces in Vision-Language ModelsJames Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou et al.NeurIPS 2023 · 13 citations
- Language-biased image classification: evaluation based on semantic representationsYoann Lemesle, Masataka Sawayama, Guillermo Valle Pérez, Maxime Adolphe et al.ICLR 2022 · 8 citations
- CLIP in Mirror: Disentangling text from visual images through reflectionTiancheng Wang, Yuguang Yang, Linlin Yang, Shaohui Lin et al.NeurIPS 2024 · 5 citations
- Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalGangyan Zeng, Yuan Zhang, Jin Wei, Dongbao Yang et al.ACM MM 2024 · 8 citations
- Variational Distribution Learning for Unsupervised Text-to-Image GenerationMinsoo Kang, Doyup Lee, Jiseob Kim, Saehoon Kim et al.CVPR 2023
