UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Rui Tian, Mingfei Gao, Mingze Xu, Jiaming Hu, Jiasen Lu, Zuxuan Wu, Yinfei Yang, Afshin Dehghan
摘要
We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 被引用 261 次
- Reconstruction Alignment Improves Unified Multimodal ModelsJi Xie, Trevor Darrell, Luke Zettlemoyer, XuDong WangICLR 2026 · 被引用 52 次
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn 等CVPR 2026 · 被引用 33 次
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated SupervisionZhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma 等ACL 2026 · 被引用 19 次
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic ControlDanfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li 等NeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- ImageGen-CoT: Enhancing Text-to-Image in-context Learning with Chain-of-Thought ReasoningJiaqi Liao, Zhengyuan Yang, Linjie Li, Dianqi Li 等ICCV 2025 · 被引用 4 次
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingHao Li, Changyao Tian, Jie Shao, Xizhou Zhu 等CVPR 2025
- dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language ModelsYi Xin, Siqi Luo, Tianxiang Xu, Qi Qin 等CVPR 2026 · 被引用 3 次
- Unsupervised Visual Chain-of-Thought Reasoning via Preference OptimizationKesen Zhao, Beier Zhu, Qianru Sun, Hanwang ZhangICCV 2025 · 被引用 3 次
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li 等ICLR 2026 · 被引用 55 次
