DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
renjie lu, Xulong Zhang, Xiaoyang Qu, Jianzong Wang, Shangfei Wang
Abstract
Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge lies in inductive biases induced by distinct supervision signals: generation branch prefers high-fidelity, fine-grained representations capable of reconstruction, while the understanding favours semantically discriminative embeddings that remain invariant to task-irrelevant factors. Consequently, optimizing these complementary but non-equivalent objectives within a monolithic backbone leads to mutual impairment instead of enhancement. In this paper, we first analyze the root cause of this interference in unified backbones and reveal a complementary structure in their internal representations. Motivated by the observation, we propose DIVA, a self-improved post-training framework that transforms the representation divergence into interior synergy. By explicitly factorizing the visual representation into shared and unique components based on two complementary information flow, DIVA enables both the understanding and generation branches to achieve beneficial transferring while preserving the integrity of unique information from cross-flow interference via mutual information estimation. Despite its generality, our method consistently achieves improvements across visual understanding (+7.82%) and generation (+8.46%). The official code is available at: https://anonymous.4open.science/r/DIVA-D225.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53608db7-22af-4ab3-9a06-a385cee7064bBuilds on20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Diffusion Feedback Helps CLIP See BetterWenxuan Wang, Quan Sun, Fan Zhang, Yepeng Tang et al.ICLR 2025
- UniGame: Turning a Unified Multimodal Model Into Its Own AdversaryZhaolong Su, Wang Lu, Hao Chen, Sharon Li et al.CVPR 2026 · 11 citations
- Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation GenerationZihan Su, Hongyang Wei, Kangrui Cen, Yong Wang et al.ICML 2026 · 15 citations
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-rewardsJixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu et al.KDD 2026 · 4 citations
- Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal ModelsJiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang et al.CVPR 2026 · 5 citations
