Reconciling Visual Perception and Generation in Diffusion Models
Liulei Li, Yi Yang, Wenguan Wang
Abstract
We present GENREP, a unified image understanding and synthesis model that jointly conducts discriminative learning and generative modeling in one training session. By leveraging Monte Carlo approximation, GENREP distills distributional knowledge embedded in diffusion models to guide the discriminative learning for visual perception tasks. Simultaneously, a semantic-driven image generation process is established, where high-level semantics learned from perception tasks can be used to inform image synthesis, creating a positive feedback loop for mutual boosts. Moreover, to reconcile the learning process for both tasks, a gradient alignment strategy is proposed to symmetrically modify the optimization directions of perception and generation losses. These designs empower GENREP to be a versatile and powerful model that achieves top-leading performance on both image understanding and generation benchmarks. Our code is available at GENREP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31a36f79-da91-41f9-8b79-d9135b67e844Cited by top-tier papers1
Ask how each one uses itBuilds on72
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image GenerationYuxin Song, Yu Lu, Haoyuan Sun, Huanjin Yao et al.CVPR 2026 · 3 citations
- SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified FlowChaoyang Wang, Xiangtai Li, Lu Qi, Henghui Ding et al.NeurIPS 2024 · 25 citations
- Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual PerceptionZiqi Pang, Xin Xu, Yu-Xiong WangICLR 2025
- Unlocking Pre-Trained Image Backbones for Semantic Image SynthesisTariq Berrada, Jakob Verbeek, Camille Couprie, Karteek AlahariCVPR 2024
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and VerifyingChengyu Bai, Jintao Chen, Xiang Bai, Yilong Chen et al.CVPR 2026 · 8 citations
