VIGC: Visual Instruction Generation and Correction
Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, Conghui He
Abstract
The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of highquality instruction-tuning data for vision-language tasks remains a challenge. The current leading paradigm, such as LLaVA, relies on language-only GPT-4 to generate data, which requires pre-annotated image captions and detection bounding boxes, suffering from understanding image details. A practical solution to this problem would be to utilize the available multimodal large language models to generate instruction data for vision-language tasks. However, it's worth noting that the currently accessible MLLMs are not as powerful as their LLM counterparts, as they tend to produce inadequate responses and generate false information. As a solution for addressing the current issue, this paper proposes the Visual Instruction Generation and Correction (VIGC) framework that enables multimodal large language models to generate instruction-tuning data and progressively enhance its quality on-the-fly. Specifically, Visual Instruction Generation (VIG) guides the vision-language model to generate diverse instruction-tuning data. To ensure generation quality, Visual Instruction Correction (VIC) adopts an iterative update mechanism to correct any inaccuracies in data produced by VIG, effectively reducing the risk of hallucination. Leveraging the diverse, high-quality data generated by VIGC, we finetune mainstream models and validate data quality based on various evaluations. Experimental results demonstrate that VIGC not only compensates for the shortcomings of language-only data generation methods, but also effectively enhances the benchmark performance. The models, datasets, and code are available at https://opendatalab.github.io/VIGC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human FeedbackTianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He et al.CVPR 2024 · 72 citations
- HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction DataQifan Yu, Juncheng Li, Longhui Wei, Liang Pang et al.CVPR 2024 · 29 citations
- DOPRA: Decoding Over-accumulation Penalization and Re-allocation in Specific Weighting LayerJinfeng Wei, Xiaofeng ZhangACM MM 2024 · 25 citations
- Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsChaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye et al.ACM MM 2024 · 19 citations
- Unleashing Region Understanding in Intermediate Layers for MLLM-based Referring Expression GenerationYaoyuan Liang, Zhuojun Cai, Jian Xu, Guanbo Huang et al.NeurIPS 2024 · 9 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Towards Self-Refinement of Vision-Language Models with Triangular ConsistencyYunlong Deng, Guangyi Chen, Tianpei Gu, Lingjing Kong et al.NeurIPS 2025 · 3 citations
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMsZiyun Dai, Xiaoqiang Li, Shaohua Zhang, Yuanchen Wu et al.ACM MM 2025
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
- Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language ReasoningRongjie Li, Yu Wu, Xuming HeCVPR 2024
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang et al.ICLR 2024 · 476 citations
