MMIFEvol: Towards Evolutionary Multimodal Instruction Following
Haoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen, Xiaojun Meng, Jiansheng Wei, Yitong Wang, Yanghua Xiao
Abstract
Multimodal Instruction Following serves as a fundamental capability of multimodal language models, involving accurate comprehension and execution of user-provided instructions. However, existing multimodal instruction-following datasets and benchmarks face the shortcomings outlined below: (a) Lack of Difficulty Stratification, they collect diverse instruction categories but neglect the stratification of difficulty levels across these categories, which leads to overlap, bias, and low interpretability. (b) Lack of Fine-Grained Metrics, they conflate the model's ability to "solve tasks" and "follow constraints" into a single metric, which fails to accurately reflect its instruction-following capability. (c) Lack of Multi-Task Instructions, they overlook the fact that real-world user instructions often consist of multiple combined tasks. This paper proposes MMIFEvol, a framework for multimodal instruction evolving and benchmarking. First, we define the essential components of a carefully curated multimodal instruction set and establish corresponding difficulty levels, based on which we synthesize diverse instruction data. Next, we decouple the evaluation criteria for the instruction following into three different metrics to construct a high-quality benchmark and assess existing models. Experimental results demonstrate that current models still struggle with following complex instructions, while fine-tuning using MMIFEvol data effectively improves models' responsiveness to multimodal instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi et al.ICLR 2024 · 315 citations
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li et al.ICCV 2025 · 37 citations
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong et al.ACL 2024 · 10 citations
- MM-IFEngine: Towards Multimodal Instruction FollowingShengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang et al.ICCV 2025 · 3 citations
Related papers
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu et al.ACL 2025 · 5 citations
- Can Large Language Models Understand Real-World Complex Instructions?Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen et al.AAAI 2024 · 99 citations
- MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMsYusu Qian, Hanrong Ye, Jean-Philippe Fauconnier, Peter Grasch et al.ICLR 2025
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling et al.ACL 2026 · 3 citations
