You Only Sample Once: Taming One-Step Text-to-Image Synthesis by Self-Cooperative Diffusion GANs
Yihong Luo, Xiaolong Chen, Xinghua Qu, Tianyang Hu, Jing Tang
Abstract
Recently, some works have tried to combine diffusion and Generative Adversarial Networks (GANs) to alleviate the computational cost of the iterative denoising inference in Diffusion Models (DMs). However, existing works in this line suffer from either training instability and mode collapse or subpar one-step generation learning efficiency. To address these issues, we introduce YOSO, a novel generative model designed for rapid, scalable, and high-fidelity one-step image synthesis with high training stability and mode coverage. Specifically, we smooth the adversarial divergence by the denoising generator itself, performing self-cooperative learning. We show that our method can serve as a one-step generation model training from scratch with competitive performance. Moreover, we extend our YOSO to one-step text-to-image generation based on pre-trained models by several effective training techniques (i.e., latent perceptual loss and latent discriminator for efficient training along with the latent DMs; the informative prior initialization (IPI), and the quick adaption stage for fixing the flawed noise scheduler). Experimental results show that YOSO achieves the state-of-the-art one-step generation performance even with Low-Rank Adaptation (LoRA) fine-tuning. In particular, we show that the YOSO-PixArt-α can generate images in one step trained on 512 resolution, with the capability of adapting to 1024 resolution without extra explicit training, requiring only ˜10 A800 days for fine-tuning. Our code is available at: https://github.com/Luo-Yihong/YOSO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbb35e50-6dc4-46d6-91db-3420e979f596Cited by top-tier papers21
- SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-TrainingJianyi Wang, Shanchuan Lin, Zhijie Lin, Yuxi Ren et al.ICLR 2026 · 51 citations
- SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image DistillationXingtong Ge, Xin Zhang, Tongda Xu, Yi Zhang et al.ICLR 2026 · 29 citations
- Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image GenerationYihong Luo, Tianyang Hu, Weijian Luo, Kenji Kawaguchi et al.NeurIPS 2025 · 20 citations
- Ultra-Fast Language Generation via Discrete Diffusion Divergence InstructHaoyang Zheng, Xinyang Liu, Cindy Xiangrui Kong, Nan Jiang et al.ICLR 2026 · 14 citations
- Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step GenerationYuyang You, Yongzhi Li, Jiahui Li, Yadong Mu et al.CVPR 2026 · 7 citations
Builds on42
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANsYanwu Xu, Yang Zhao, Zhisheng Xiao, Tingbo HouCVPR 2024
- Generative Adversarial DiffusionU-Chae Jun, Jaeeun Ko, Jiwoo KangICCV 2025 · 2 citations
- Revisiting Diffusion Models: From Generative Pre-training to One-Step GenerationBowen Zheng, Tianming YangICML 2025
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 319 citations
- SF-V: Single Forward Video Generation ModelZhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu et al.NeurIPS 2024 · 43 citations
