Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion
Xuantong Liu, Tianyang Hu, Wenjia Wang, Kenji Kawaguchi, Yuan Yao
Abstract
As a dominant force in text-to-image generation tasks, Diffusion Probabilistic Models (DPMs) face a critical challenge in controllability, struggling to adhere strictly to complex, multi-faceted instructions. In this work, we aim to address this alignment challenge for conditional generation tasks. First, we provide an alternative view of state-of-the-art DPMs as a way of inverting advanced Vision-Language Models (VLMs). With this formulation, we naturally propose a training-free approach that bypasses the conventional sampling process associated with DPMs. By directly optimizing images with the supervision of discriminative VLMs, the proposed method can potentially achieve a better text-image alignment. As proof of concept, we demonstrate the pipeline with the pre-trained BLIP-2 model and identify several key designs for improved image generation. To further enhance the image fidelity, a Score Distillation Sampling module of Stable Diffusion is incorporated. By carefully balancing the two components during optimization, our method can produce high-quality images with near state-of-the-art performance on T2I-Compbench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3548efb6-0ee7-4a03-bc23-e97a5dd0fe5cCited by top-tier papers5
- Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image GenerationYihong Luo, Tianyang Hu, Weijian Luo, Kenji Kawaguchi et al.NeurIPS 2025 · 20 citations
- TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable RewardYihong Luo, Tianyang Hu, Weijian Luo, Jing TangICML 2026 · 5 citations
- GENMAC: Compositional Text-to-Video Generation with Multi-Agent CollaborationKaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin et al.AAAI 2026 · 1 citation
- T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video GenerationKaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu et al.CVPR 2025
- Open-Vocabulary Customization from CLIP via Data-Free Knowledge DistillationYongxian Wei, Zixuan Hu, Li Shen, Zhenyi Wang et al.ICLR 2025
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination CapabilityRunhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang et al.ICCV 2023 · 10 citations
- Training-Free Diffusion Model Alignment with Sampling DemonsPo-Hung Yeh, Kuang-Huei Lee, Jun-Cheng ChenICLR 2025
- Free2 Guide: Training-Free Text-to-Video Alignment Using Image LVLMJaemin Kim, Bryan Sangwoo Kim, Jong Chul YeICCV 2025 · 1 citation
- Are Diffusion Models Vision-And-Language Reasoners?Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Chris Pal et al.NeurIPS 2023 · 21 citations
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image GenerationPaul Grimal, Michaël Soumm, Hervé Le Borgne, Olivier Ferret et al.AAAI 2026 · 1 citation
