MM-IFEngine: Towards Multimodal Instruction Following
Shengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang, Haodong Duan, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Dahua Lin, Jiaqi Wang
Abstract
The Instruction Following (IF) ability measures how well Multi-modal Large Language Models (MLLMs) understand exactly what users are telling them and whether they are doing it right. Existing multimodal instruction following training data is scarce, the benchmarks are simple with atomic instructions, and the evaluation strategies are imprecise for tasks demanding exact output constraints. To address this, we present MM-IFEngine, an effective pipeline to generate high-quality image-instruction pairs. Our MM-IFEngine pipeline yields large-scale, diverse, and high-quality training data MM-IFInstruct-23k, which is suitable for Supervised Fine-Tuning (SFT) and extended as MM-IFDPO-23k for Direct Preference Optimization (DPO). We further introduce MM-IFEval, a challenging and diverse multi-modal instruction-following benchmark that includes (1) both compose-level constraints for output responses and perception-level constraints tied to the input images, and (2) a comprehensive evaluation pipeline incorporating both rule-based assessment and judge model. We conduct SFT and DPO experiments and demonstrate that fine-tuning MLLMs on MM-IFInstruct-23k and MM-IFDPO-23k achieves notable gains on various IF benchmarks, such as MM-IFEval (+10.2), MIA (+7.6), and IFEval (+12.3). We have fully open-sourced the datasets (both SFT and DPO), evaluation code and training scripts at https://github.com/SYuan03/MM-IFEngine.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7205a4c0-48d6-47a6-9cb8-d3d772b3c5f5Cited by top-tier papers14
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang et al.CVPR 2026 · 179 citations
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from ExperienceZEYI SUN, Ziyu Liu, Yuhang Zang, Yuhang Cao et al.ICML 2026 · 58 citations
- SIM-CoT: Supervised Implicit Chain-of-ThoughtXilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong et al.ICLR 2026 · 58 citations
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun et al.NeurIPS 2025 · 17 citations
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He et al.ICLR 2026 · 17 citations
Builds on22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- MMIFEvol: Towards Evolutionary Multimodal Instruction FollowingHaoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen et al.AAAI 2026 · 1 citation
- MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMsYusu Qian, Hanrong Ye, Jean-Philippe Fauconnier, Peter Grasch et al.ICLR 2025
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke et al.ACL 2026 · 2 citations
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling et al.ACL 2026 · 3 citations
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng et al.ACL 2025 · 3 citations
