MM-IFEngine: Towards Multimodal Instruction Following
Shengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang, Haodong Duan, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Dahua Lin, Jiaqi Wang
摘要
The Instruction Following (IF) ability measures how well Multi-modal Large Language Models (MLLMs) understand exactly what users are telling them and whether they are doing it right. Existing multimodal instruction following training data is scarce, the benchmarks are simple with atomic instructions, and the evaluation strategies are imprecise for tasks demanding exact output constraints. To address this, we present MM-IFEngine, an effective pipeline to generate high-quality image-instruction pairs. Our MM-IFEngine pipeline yields large-scale, diverse, and high-quality training data MM-IFInstruct-23k, which is suitable for Supervised Fine-Tuning (SFT) and extended as MM-IFDPO-23k for Direct Preference Optimization (DPO). We further introduce MM-IFEval, a challenging and diverse multi-modal instruction-following benchmark that includes (1) both compose-level constraints for output responses and perception-level constraints tied to the input images, and (2) a comprehensive evaluation pipeline incorporating both rule-based assessment and judge model. We conduct SFT and DPO experiments and demonstrate that fine-tuning MLLMs on MM-IFInstruct-23k and MM-IFDPO-23k achieves notable gains on various IF benchmarks, such as MM-IFEval (+10.2), MIA (+7.6), and IFEval (+12.3). We have fully open-sourced the datasets (both SFT and DPO), evaluation code and training scripts at https://github.com/SYuan03/MM-IFEngine.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang 等CVPR 2026 · 被引用 179 次
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from ExperienceZEYI SUN, Ziyu Liu, Yuhang Zang, Yuhang Cao 等ICML 2026 · 被引用 58 次
- SIM-CoT: Supervised Implicit Chain-of-ThoughtXilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong 等ICLR 2026 · 被引用 58 次
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun 等NeurIPS 2025 · 被引用 17 次
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
相关 Paper
- MMIFEvol: Towards Evolutionary Multimodal Instruction FollowingHaoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen 等AAAI 2026 · 被引用 1 次
- MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMsYusu Qian, Hanrong Ye, Jean-Philippe Fauconnier, Peter Grasch 等ICLR 2025
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke 等ACL 2026 · 被引用 2 次
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling 等ACL 2026 · 被引用 3 次
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng 等ACL 2025 · 被引用 3 次
