Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation
Zihan Su, Hongyang Wei, Kangrui Cen, Yong Wang, Guanhua CHEN, Chun Yuan, Xiangxiang Chu
摘要
Unified Multimodal Models (UMMs) integrate both visual understanding and generation within a single framework. Their ultimate aspiration is to create a cycle where understanding and generation mutually reinforce each other. While recent post-training methods have successfully leveraged understanding to enhance generation, the reverse direction of utilizing generation to improve understanding remains largely unexplored. In this work, we propose UniMRG (Unified Multi-Representation Generation), a simple yet effective architecture-agnostic post-training method. UniMRG enhances the understanding capabilities of UMMs by incorporating auxiliary generation tasks. Specifically, we train UMMs to generate multiple intrinsic representations of input images, namely pixel (reconstruction), depth (geometry), and segmentation (structure), alongside standard visual understanding objectives. By synthesizing these diverse representations, UMMs capture complementary information regarding appearance, spatial relations, and structural layout. Consequently, UMMs develop a deeper and more comprehensive understanding of visual inputs. Extensive experiments across diverse UMM architectures demonstrate that our method notably enhances fine-grained perception, reduces hallucinations, and improves spatial understanding, while simultaneously boosting generation capabilities. †Work done during internship at AMAP, Alibaba Group * Equal contribution ‡Project lead
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model MergingYongxian Wei, Runxi Cheng, Weike Jin, Enneng Yang 等ICLR 2026 · 被引用 10 次
- EasyTune: Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion GenerationXiaofeng Tan, Wanjiang Weng, Haodong Lei, Hongsong WangICLR 2026 · 被引用 6 次
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask PredictionShannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
相关 Paper
- DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcementrenjie lu, Xulong Zhang, Xiaoyang Qu, Jianzong Wang 等ICML 2026
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 等CVPR 2026 · 被引用 36 次
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodingYueming Xu, Jiahui Zhang, Ze Huang, Yurui Chen 等ICLR 2026 · 被引用 8 次
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao 等ACL 2021
- Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal ModelsJiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang 等CVPR 2026 · 被引用 5 次
