Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!
Zaid Khan, B. G. Vijay Kumar, Samuel Schulter, Xiang Yu, Yun Fu, Manmohan Chandraker
摘要
Finetuning a large vision language model (VLM) on a target dataset after large scale pretraining is a dominant paradigm in visual question answering (VQA). Datasets for specialized tasks such as knowledge-based VQA or VQA in non natural-image domains are orders of magnitude smaller than those for general-purpose VQA. While collecting additional labels for specialized tasks or domains can be challenging, unlabeled images are often available. We introduce SelTDA (Self-Taught Data Augmentation), a strategy for finetuning large VLMs on small-scale VQA datasets. SelTDA uses the VLM and target dataset to build a teacher model that can generate question-answer pseudolabels directly conditioned on an image alone, allowing us to pseudolabel unlabeled images. SelTDA then finetunes the initial VLM on the original dataset augmented with freshly pseudolabeled images. We describe a series of experiments showing that our self-taught data augmentation increases robustness to adversarially searched questions, counterfactual examples and rephrasings, improves domain generalization, and results in greater retention of numerical reasoning skills. The proposed strategy requires no additional annotations or architectural modifications, and is compatible with any modern encoder-decoder multimodal transformer. Code available at https://github.com/codezakh/SelTDA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Exploring Question Decomposition for Zero-Shot VQAZaid Khan, Vijay Kumar B. G, Samuel Schulter, Manmohan Chandraker 等NeurIPS 2023 · 被引用 26 次
- CHAIN: Enhancing Generalization in Data-Efficient GANs via LipsCHitz Continuity ConstrAIned NormalizationYao Ni, Piotr KoniuszCVPR 2024 · 被引用 10 次
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceHaicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju 等ICCV 2025 · 被引用 3 次
- Leveraging Textual Compositional Reasoning for Robust Change CaptioningKyu Ri Park, Jiyoung Park, Seong Tae Kim, Hong Joo Lee 等AAAI 2026
- Self-Training Large Language Models for Improved Visual Program Synthesis With Visual ReinforcementZaid Khan, Vijay Kumar B. G, Samuel Schulter, Yun Fu 等CVPR 2024
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Video-STaR: Self-Training Enables Video Instruction Tuning with Any SupervisionOrr Zohar, Xiaohan Wang, Yonatan Bitton, Idan Szpektor 等ICLR 2025
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu 等ACL 2022
- Discovering the Unknown Knowns: Turning Implicit Knowledge in the Dataset into Explicit Training Examples for Visual Question AnsweringJihyung Kil, Cheng Zhang, Dong Xuan, Wei-Lun ChaoEMNLP 2021 · 被引用 17 次
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti 等NeurIPS 2024 · 被引用 20 次
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingGi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak 等CVPR 2023
