Bootstrapping Vision-Language Learning with Decoupled Language Pre-training
Yiren Jian, Chongyang Gao, Soroush Vosoughi
摘要
We present a novel methodology aimed at optimizing the application of frozen large language models (LLMs) for resource-intensive vision-language (VL) pre-training. The current paradigm uses visual features as prompts to guide language models, with a focus on determining the most relevant visual features for corresponding text. Our approach diverges by concentrating on the language component, specifically identifying the optimal prompts to align with visual features. We introduce the Prompt-Transformer (P-Former), a model that predicts these ideal prompts, which is trained exclusively on linguistic data, bypassing the need for image-text pairings. This strategy subtly bifurcates the end-to-end VL training process into an additional, separate stage. Our experiments reveal that our framework significantly enhances the performance of a robust image-to-text baseline (BLIP-2), and effectively narrows the performance gap between models trained with either 4M or 129M image-text pairs. Importantly, our framework is modality-agnostic and flexible in terms of architectural design, as validated by its successful application in a video learning task using varied base modules. The code will be made available at https://github.com/yiren-jian/BLIText .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought VerificationLinzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu 等ACL 2025 · 被引用 38 次
- Towards Neuron Attributions in Multi-Modal Large Language ModelsJunfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang 等NeurIPS 2024 · 被引用 16 次
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin 等NeurIPS 2025 · 被引用 11 次
- Growing Through Experience: Scaling Episodic Grounding in Language ModelsChunhui Zhang, Sirui Wang, Zhongyu Ouyang, Xiangchi Yuan 等ACL 2025 · 被引用 6 次
- Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention ReasoningYu Wang, Shengjie ZhaoCVPR 2026 · 被引用 6 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionYiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang 等ACL 2024 · 被引用 5 次
- UP-DP: Unsupervised Prompt Learning for Data Pre-Selection with Vision-Language ModelsXin Li, Sima Behpour, Thang Long Doan, Wenbin He 等NeurIPS 2023 · 被引用 4 次
- Frozen Transformers in Language Models Are Effective Visual Encoder LayersZiqi Pang, Ziyang Xie, Yunze Man, Yu-Xiong WangICLR 2024 · 被引用 54 次
- A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language ModelsWoojeong Jin, Yu Cheng, Yelong Shen, Weizhu Chen 等ACL 2022
- Advancing Prompt Learning through an External LayerFangming Cui, Xun Yang, Chao Wu, Liang Xiao 等ACM MM 2024 · 被引用 3 次
