SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-rewards
Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu, Ji-Rong Wen, Rui Yan
摘要
Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-modal understanding and generation into a single framework. However, LMMs still struggle to achieve accurate vision-language alignment, prone to generating text responses contradicting the visual input or failing to follow the text-to-image prompts. To alleviate these issues, a promising line of research explores improving LMMs in the post-training stage. However, most existing solutions rely on external supervision (e.g., human annotations, reward models) and are typically tailored to only unidirectional tasks, i.e., optimizing either vision understanding or generation. In this work, based on the observation that understanding and generation are naturally inverse dual tasks, we propose SUDER (Self-improving Unified LMMs with Dual sElf-Rewards), a framework reinforcing the understanding and generation capabilities of LMMs with a self-supervised dual reward mechanism. SUDER leverages the inherent duality between understanding and generation to provide self-supervised optimization signals for each other. Specifically, we sample multiple outputs for a given input, then reverse the input-output pairs to compute the dual likelihood within the model as self-rewards for optimization. Extensive experimental results on visual understanding and generation benchmarks demonstrate that our method can effectively enhance the performance of the LMM without any external supervision, especially achieving remarkable improvements in text-to-image tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal ModelsJiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang 等CVPR 2026 · 被引用 5 次
- Self-Corrected Image Generation with Explainable Latent RewardsYinyi Luo, Hrishikesh Gokhale, Marios Savvides, Jindong Wang 等CVPR 2026 · 被引用 1 次
- STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMsZongzhao Li, Zongyang Ma, Mingze Li, Songyou Li 等CVPR 2026
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image GenerationYoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi 等CVPR 2026
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang 等NeurIPS 2025 · 被引用 15 次
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhanceChunwei Wang, Guansong Lu, Junwei Yang, Runhui Huang 等ICCV 2025 · 被引用 5 次
- X-Fusion: Introducing New Modality to Frozen Large Language ModelsSicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer 等ICCV 2025
- Calibrated Self-Rewarding Vision Language ModelsYiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang 等NeurIPS 2024 · 被引用 77 次
- UnifiedVisual: A Framework for Constructing Unified Vision-Language DatasetsPengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang 等EMNLP 2025
