CLIPPO: Image-and-Language Understanding from Pixels Only
Michael Tschannen, Basil Mustafa, Neil Houlsby
摘要
Multimodal models are becoming increasingly effective, in part due to unified components, such as the Transformer architecture. However, multimodal models still often consist of many task-and modality-specific pieces and training procedures. For example, CLIP (Radford et al., 2021) trains independent text and image towers via a contrastive loss. We explore an additional unification: the use of a pure pixelbased model to perform image, text, and multimodal tasks. Our model is trained with contrastive loss alone, so we call it CLIP-Pixels Only (CLIPPO). CLIPPO uses a single encoder that processes both regular images and text rendered as images. CLIPPO performs image-based tasks such as retrieval and zero-shot image classification almost as well as CLIP-style models, with half the number of parameters and no text-specific tower or embedding. When trained jointly via image-text contrastive learning and next-sentence contrastive learning, CLIPPO can perform well on natural language understanding tasks, without any word-level loss (language modelling or masked language modelling), outperforming pixel-based prior work. Surprisingly, CLIPPO can obtain good accuracy in visual question answering, simply by rendering the question and image together. Finally, we exploit the fact that CLIPPO does not require a tokenizer to show that it can achieve strong performance on multilingual multimodal retrieval without modifications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Image Captioners Are Scalable Vision Learners TooMichael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai 等NeurIPS 2023 · 被引用 104 次
- LoGoPrompt: Synthetic Text Images Can Be Good Visual Prompts for Vision-Language ModelsCheng Shi, Sibei YangICCV 2023 · 被引用 35 次
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu 等NeurIPS 2025 · 被引用 32 次
- Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal LearningAlex Jinpeng Wang, Linjie Li, Yiqi Lin, Min Li 等NeurIPS 2024 · 被引用 21 次
- Cross-Lingual Text-Rich Visual Comprehension: An Information Theory PerspectiveXinmiao Yu, Xiaocheng Feng, Yun Li, Minghui Liao 等AAAI 2025 · 被引用 7 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene DescriptionsIoanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat 等EMNLP 2025 · 被引用 1 次
- VISTA: Visualized Text Embedding For Universal Multi-Modal RetrievalJunjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao 等ACL 2024
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 被引用 33 次
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 被引用 21 次
