RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Yang Shi, Yuhao Dong, Yue Ding, Yuran Wang, Xuanyu Zhu, Sheng Zhou, Wenting Liu, Haochen Tian, Rundong Wang, Huanqian Wang, Zuyan Liu, Bohan Zeng
Abstract
The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: does this architectural unification actually enable synergetic interaction between the constituent capabilities? Existing evaluation paradigms, which primarily assess understanding and generation in isolation, are insufficient for determining whether a unified model can leverage its understanding to enhance its generation, or use generative simulation to facilitate deeper comprehension. To address this critical gap, we introduce RealUnify, a benchmark specifically designed to evaluate bidirectional capability synergy. RealUnify comprises 1,000 meticulously human-annotated instances spanning 10 categories and 32 subtasks. It is structured around two core axes: 1) Understanding Enhances Generation, which requires reasoning (e.g., commonsense, logic) to guide image generation, and 2) Generation Enhances Understanding, which necessitates mental simulation or reconstruction (e.g., of transformed or disordered visual inputs) to solve reasoning tasks. A key contribution is our dual-evaluation protocol, which combines direct end-to-end assessment with a diagnostic stepwise evaluation that decomposes tasks into distinct understanding and generation phases. This protocol allows us to precisely discern whether performance bottlenecks stem from deficiencies in core abilities or from a failure to integrate them. Through large-scale evaluations of 12 leading unified models and 6 specialized baselines, we find that current unified models still struggle to achieve effective synergy, indicating that architectural unification alone is insufficient. These results highlight the need for new training strategies and inductive biases to fully unlock the potential of unified modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f705cfbf-628f-4044-84e8-9eeb294baac5Cited by top-tier papers6
- SimScale: Learning to Drive via Real-World Simulation at ScaleHaochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang et al.CVPR 2026 · 40 citations
- FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis NetworkFangtong Sun, Congyu Li, Ke Yang, Yuchen Pan et al.NeurIPS 2025 · 4 citations
- UniVerse: Empower Unified Generation with Reasoning and KnowledgeKaiyue Sun, Weiyang Jin, Chengqi Duan, Rongyao Fang et al.CVPR 2026
- Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality ConflictsZhuoran Zhang, Tengyue Wang, Xilin Gong, Yang Shi et al.ICML 2026
- Monet: Reasoning in Latent Visual Space Beyond Image and LanguageQixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang et al.CVPR 2026
Builds on13
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan et al.CVPR 2026 · 231 citations
- WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationYuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin et al.ICML 2026 · 195 citations
- Thyme: Think Beyond ImagesYifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu et al.ICLR 2026 · 146 citations
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen et al.CVPR 2026 · 124 citations
Related papers
- Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian et al.ACL 2026 · 19 citations
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu et al.ICLR 2026 · 15 citations
- Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and GenerationJinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang JiangICML 2026 · 2 citations
- MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation ModelsWulin Xie, YiFan Zhang, Chaoyou Fu, Yang Shi et al.ICLR 2026 · 31 citations
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal GenerationYongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma et al.ICLR 2026 · 13 citations
