HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
Ling Yang, Xinchen Zhang, Ye Tian, Shiyi Zhang, Chenming Shang, Minghao Xu, Wentao Zhang, Bin Cui
Abstract
The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with powerful models like Show-o, Transfusion and Emu3 achieving notable progress in unified image understanding and generation. For the first time, we uncover a common phenomenon: the understanding capabilities of MLLMs are typically stronger than their generative capabilities, with a significant gap between the two. Building on this insight, we propose HermesFlow, a simple yet general framework designed to seamlessly bridge the gap between understanding and generation in MLLMs. Specifically, we take the homologous data as input to curate homologous preference data of both understanding and generation. Through Pair-DPO and self-play iterative optimization, HermesFlow effectively aligns multimodal understanding and generation using homologous preference data. Extensive experiments demonstrate the significant superiority of our approach over prior methods, particularly in narrowing the gap between multimodal understanding and generation. These findings highlight the potential of HermesFlow as a general alignment framework for next-generation multimodal foundation models.
Recently, there has been growing interest in exploring the synergy between multimodal understanding and generation [45,40,3]. Liquid [45] demonstrates that these two tasks are mutually * Contributed equally.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a524e63-aeab-4ead-98f7-e0fbbca07e26Cited by top-tier papers4
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang et al.NeurIPS 2025 · 15 citations
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language ReasoningYizhen Zhang, Yang Ding, Shuoshuo Zhang, Xinchen Zhang et al.NeurIPS 2025 · 13 citations
- Multimodal Meta-Verifier with Explicit Structured RecalibrationXinchen Zhang, Bowei Liu, Jiale Liu, Chufan Shi et al.ICML 2026 · 1 citation
- Self-Corrected Image Generation with Explainable Latent RewardsYinyi Luo, Hrishikesh Gokhale, Marios Savvides, Jindong Wang et al.CVPR 2026 · 1 citation
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMsYujin Han, Hao Chen, Andi Han, Zhiheng Wang et al.ICLR 2026 · 9 citations
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingHao Li, Changyao Tian, Jie Shao, Xizhou Zhu et al.CVPR 2025
- Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal ModelsJiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang et al.CVPR 2026 · 5 citations
- Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal ModelsSen Ye, Mengde Xu, Shuyang Gu, Di He et al.ICLR 2026 · 5 citations
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu et al.NeurIPS 2025 · 3 citations
