Disentangling visual and written concepts in CLIP
Joanna Materzynska, Antonio Torralba, David Bau
摘要
The CLIP network measures the similarity between natural text and images; in this work, we investigate the entanglement of the representation of word images and natural images in its image encoder. First, we find that the image encoder has an ability to match word images with natural images of scenes described by those words. This is consistent with previous research that suggests that the meaning and the spelling of a word might be entangled deep within the network. On the other hand, we also find that CLIP has a strong ability to match nonsense words, suggesting that processing of letters is separated from processing of their meaning. To explicitly determine whether the spelling capability of CLIP is separable, we devise a procedure for identifying representation subspaces that selectively isolate or eliminate spelling capabilities. We benchmark our methods against a range of retrieval tasks, and we also test them by measuring the appearance of text in CLIP-guided generated images. We find that our methods are able to cleanly separate spelling capabilities of CLIP from the visual processing of natural images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 被引用 262 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- Interpreting CLIP's Image Representation via Text-Based DecompositionYossi Gandelsman, Alexei A. Efros, Jacob SteinhardtICLR 2024 · 被引用 179 次
- What the DAAM: Interpreting Stable Diffusion Using Cross AttentionRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang 等ACL 2023 · 被引用 93 次
- Decomposing and Editing Predictions by Modeling Model ComputationHarshay Shah, Andrew Ilyas, Aleksander MadryICML 2024 · 被引用 25 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 被引用 1,049 次
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 被引用 421 次
相关 Paper
- Parts of Speech-Grounded Subspaces in Vision-Language ModelsJames Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou 等NeurIPS 2023 · 被引用 13 次
- Language-biased image classification: evaluation based on semantic representationsYoann Lemesle, Masataka Sawayama, Guillermo Valle Pérez, Maxime Adolphe 等ICLR 2022 · 被引用 8 次
- CLIP in Mirror: Disentangling text from visual images through reflectionTiancheng Wang, Yuguang Yang, Linlin Yang, Shaohui Lin 等NeurIPS 2024 · 被引用 5 次
- Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalGangyan Zeng, Yuan Zhang, Jin Wei, Dongbao Yang 等ACM MM 2024 · 被引用 8 次
- Variational Distribution Learning for Unsupervised Text-to-Image GenerationMinsoo Kang, Doyup Lee, Jiseob Kim, Saehoon Kim 等CVPR 2023
