MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
Wulin Xie, YiFan Zhang, Chaoyou Fu, Yang Shi, Jianshu Zeng, Bingyan Nie, Hongkai Chen, Zhang Zhang, Liang Wang
摘要
Unified Multimodal Large Language Models (U-MLLMs) have garnered considerable interest for their ability to seamlessly integrate generation and comprehension tasks. However, existing research lacks a unified evaluation standard, often relying on isolated benchmarks to assess these capabilities. Moreover, current work highlights the potential of “mixed-modality generation capabilities” through case studies—such as generating auxiliary lines in images to solve geometric problems, or reasoning through a problem before generating a corresponding image. Despite this, there is no standardized benchmark to assess models on such unified tasks. To address this gap, we introduce MME-Unify, also termed as MME-U, the first open and reproducible benchmark designed to evaluate multimodal comprehension, generation, and mixed-modality generation capabilities. For comprehension and generation tasks, we curate a diverse set of tasks from 12 datasets, aligning their formats and metrics to develop a standardized evaluation framework. For unified tasks, we design five subtasks to rigorously assess how models’ understanding and generation capabilities can mutually enhance each other. Evaluation of 17 U-MLLMs, including Janus-Pro, Bagel, and Gemini2-Flash, reveals significant room for improvement, particularly in areas such as instruction following and image generation quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement LearningYifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu 等ICLR 2026 · 被引用 65 次
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive BenchmarkYang Shi, Yuhao Dong, Yue Ding, Yuran Wang 等CVPR 2026 · 被引用 35 次
- Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian 等ACL 2026 · 被引用 19 次
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and GenerationLing Yang, Xinchen Zhang, Ye Tian, Shiyi Zhang 等NeurIPS 2025 · 被引用 16 次
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu 等CVPR 2026 · 被引用 7 次
它引用的顶会 Paper22
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 被引用 403 次
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang 等NeurIPS 2025 · 被引用 234 次
- Making LLaMA SEE and Draw with SEED TokenizerYuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge 等ICLR 2024 · 被引用 202 次
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua 等NeurIPS 2024 · 被引用 100 次
相关 Paper
- Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and GenerationJinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang JiangICML 2026 · 被引用 2 次
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang 等CVPR 2026 · 被引用 10 次
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu 等ICLR 2026 · 被引用 15 次
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsPeng Xia, Siwei Han, Shi Qiu, Yiyang Zhou 等ICLR 2025
- OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li 等CVPR 2025
