Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution
Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding
Abstract
Human preference alignment can significantly enhance the capabilities of Multimodal Large Language Models (MLLMs). However, collecting high-quality preference data remains costly. One promising solution is the self-evolution strategy, where models are iteratively trained on data they generate. Current multimodal self-evolution techniques, nevertheless, still need human-or GPT-annotated data. Some methods even require extra models or ground truth answers to construct preference data. To overcome these limitations, we propose a novel multimodal self-evolution framework that empowers the model to autonomously generate high-quality questions and answers using only unannotated images. First, in the question generation phase, we implement an imagedriven self-questioning mechanism. This approach allows the model to create questions and evaluate their relevance and answerability based on the image content. If a question is deemed irrelevant or unanswerable, the model regenerates it to ensure alignment with the image. This process establishes a solid foundation for subsequent answer generation and optimization. Second, while generating answers, we design an answer self-enhancement technique to boost the discriminative power of answers. We begin by captioning the images and then use the descriptions to enhance the generated answers. Additionally, we utilize corrupted images to generate rejected answers, thereby forming distinct preference pairs for effective optimization. Finally, in the optimization step, we incorporate an image content alignment loss function alongside the Direct Preference Optimization (DPO) loss to mitigate hallucinations. This function maximizes the likelihood of the above generated descriptions in order to constrain the model's attention to the image content. As a result, model can generate more accurate and reliable outputs. Experiments demonstrate that our framework is competitively compared with previous methods that utilize external information, paving the way for more efficient and scalable MLLMs. The code is available at https://github.com/WentaoTan/SENA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 615cb24f-406b-4f1e-8484-2f4f0cf89f85Cited by top-tier papers5
- First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-TrainingLai Wei, Yuting Li, Chen Wang, Yue Wang et al.NeurIPS 2025 · 28 citations
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal UnderstandingXin Jin, Siyuan Li, Siyong Jian, Kai Yu et al.ICLR 2026 · 14 citations
- Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent RecognitionQianrui Zhou, Hua Xu, Yunjin Gu, Yifan Wang et al.CVPR 2026 · 3 citations
- Effortless Active Labeling for Long-Term Test-Time AdaptationGuowei Wang, Changxing DingCVPR 2025
- Dual-Estimator: Decoupling Global and Local Semantic Shift for Drift Compensation in Class-Incremental LearningFankang Xu, Lu Jin, Yanpeng Sun, Shiyu Xuan et al.CVPR 2026
Builds on18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image GenerationYoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi et al.CVPR 2026
- EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language ModelsShangyu Xing, Fei Zhao, Zhen Wu, Tuo An et al.EMNLP 2024 · 6 citations
- Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMsZitian Wang, Yue Liao, Kang Rong, Fengyun Rao et al.ICCV 2025
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidanceJiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang et al.ICCV 2025 · 4 citations
- Improving Large Vision and Language Models by Learning from a Panel of PeersJefferson Hernandez, Jing Shi, Simon Jenni, Vicente Ordonez et al.ICCV 2025
