Visual Instruction Tuning with Polite Flamingo
Delong Chen, Jianfeng Liu, Wenliang Dai, Baoyuan Wang
摘要
Recent research has demonstrated that the multi-task fine-tuning of multi-modal Large Language Models (LLMs) using an assortment of annotated downstream vision-language datasets significantly enhances their performance. Yet, during this process, a side effect, which we termed as the "multi-modal alignment tax", surfaces. This side effect negatively impacts the model's ability to format responses appropriately - for instance, its "politeness" - due to the overly succinct and unformatted nature of raw annotations, resulting in reduced human preference. In this paper, we introduce Polite Flamingo, a multi-modal response rewriter that transforms raw annotations into a more appealing, "polite" format. Polite Flamingo is trained to reconstruct high-quality responses from their automatically distorted counterparts and is subsequently applied to a vast array of vision-language datasets for response rewriting. After rigorous filtering, we generate the PF-1M dataset and further validate its value by fine-tuning a multi-modal LLM with it. Combined with novel methodologies including U-shaped multi-stage tuning and multi-turn augmentation, the resulting model, Clever Flamingo, demonstrates its advantages in both multi-modal understanding and response politeness according to automated and human evaluations. Code and dataset are available at https://github.com/ChenDelong1999/polite-flamingo
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Emu: Generative Pretraining in MultimodalityQuan Sun, Qiying Yu, Yufeng Cui, Fan Zhang 等ICLR 2024 · 被引用 161 次
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung 等ICLR 2026 · 被引用 60 次
- CapsFusion: Rethinking Image-Text Data at ScaleQiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui 等CVPR 2024 · 被引用 17 次
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal ModelsXiwen Wei, Mustafa Munir, Radu MarculescuNeurIPS 2025 · 被引用 9 次
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee 等EMNLP 2024 · 被引用 8 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Do DALL-E and Flamingo Understand Each Other?Hang Li, Jindong Gu, Rajat Koner, Sahand Sharifzadeh 等ICCV 2023 · 被引用 14 次
- FLAME: Learning to Navigate with Multimodal LLM in Urban EnvironmentsYunzhe Xu, Yiyuan Pan, Zhe Liu, Hesheng WangAAAI 2025 · 被引用 3 次
- Robust Multimodal Large Language Models Against Modality ConflictZongmeng Zhang, Wengang Zhou, Jie Zhao, Houqiang LiICML 2025
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?Yanbo Wang, Jiyang Guan, Jian Liang, Ran HeCVPR 2025
- ZINA: Multimodal Fine-grained Hallucination Detection and EditingYuiga Wada, Kazuki Matsuda, Komei Sugiura, Graham NeubigCVPR 2026 · 被引用 5 次
