SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-rewards
Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu, Ji-Rong Wen, Rui Yan
Abstract
Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-modal understanding and generation into a single framework. However, LMMs still struggle to achieve accurate vision-language alignment, prone to generating text responses contradicting the visual input or failing to follow the text-to-image prompts. To alleviate these issues, a promising line of research explores improving LMMs in the post-training stage. However, most existing solutions rely on external supervision (e.g., human annotations, reward models) and are typically tailored to only unidirectional tasks, i.e., optimizing either vision understanding or generation. In this work, based on the observation that understanding and generation are naturally inverse dual tasks, we propose SUDER (Self-improving Unified LMMs with Dual sElf-Rewards), a framework reinforcing the understanding and generation capabilities of LMMs with a self-supervised dual reward mechanism. SUDER leverages the inherent duality between understanding and generation to provide self-supervised optimization signals for each other. Specifically, we sample multiple outputs for a given input, then reverse the input-output pairs to compute the dual likelihood within the model as self-rewards for optimization. Extensive experimental results on visual understanding and generation benchmarks demonstrate that our method can effectively enhance the performance of the LMM without any external supervision, especially achieving remarkable improvements in text-to-image tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9aea23df-00a9-40ea-bff0-06884fbee696Cited by top-tier papers4
- Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal ModelsJiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang et al.CVPR 2026 · 5 citations
- Self-Corrected Image Generation with Explainable Latent RewardsYinyi Luo, Hrishikesh Gokhale, Marios Savvides, Jindong Wang et al.CVPR 2026 · 1 citation
- STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMsZongzhao Li, Zongyang Ma, Mingze Li, Songyou Li et al.CVPR 2026
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image GenerationYoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi et al.CVPR 2026
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang et al.NeurIPS 2025 · 15 citations
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhanceChunwei Wang, Guansong Lu, Junwei Yang, Runhui Huang et al.ICCV 2025 · 5 citations
- X-Fusion: Introducing New Modality to Frozen Large Language ModelsSicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer et al.ICCV 2025
- Calibrated Self-Rewarding Vision Language ModelsYiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang et al.NeurIPS 2024 · 77 citations
- UnifiedVisual: A Framework for Constructing Unified Vision-Language DatasetsPengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang et al.EMNLP 2025
