MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
Hui Li, Jiayue Lyu, Fu-Yun Wang, Kaihui Cheng, Siyu Zhu, Jingdong Wang
Abstract
This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at the training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and during testing, the input is the generated noisy data. We present a novel training approach, named MixFlow, for improving the training performance. Our approach is motivated by the Slow Flow phenomenon: the ground-truth interpolation that is the nearest to the generated noisy data at a given sampling timestep is observed to correspond to a higher-noise timestep (termed slowed timestep), i.e., the corresponding ground-truth timestep is slower than the sampling timestep. MixFlow leverages the interpolations at the slowed timesteps, named slowed interpolation mixture, for post-training the prediction network at each training timestep. Experiments over class-conditional image generation (including SiT, REPA, and RAE) and text-to-image generation, validate the effectiveness of our approach. Our approach MixFlow over the RAE models achieve strong generation results on ImageNet: FID (without guidance) and (with guidance) at , and FID (without guidance) and (with guidance) at .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb53c15d-5a03-4bec-adc3-4659406fdd1dBuilds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Input Perturbation Reduces Exposure Bias in Diffusion ModelsMang Ning, Enver Sangineto, Angelo Porrello, Simone Calderara et al.ICML 2023 · 100 citations
- Anti-Exposure Bias in Diffusion ModelsJunyu Zhang, Daochang Liu, Eunbyung Park, Shichao Zhang et al.ICLR 2025
- AID: Attention Interpolation of Text-to-Image DiffusionQiyuan He, Jinghao Wang, Ziwei Liu, Angela YaoNeurIPS 2024 · 30 citations
- Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion ModelsZhiyao Ren, Yibing Zhan, Liang Ding, Gaoang Wang et al.AAAI 2024 · 15 citations
- Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion ModelMin Zhao, Hongzhou Zhu, Chendong Xiang, Kaiwen Zheng et al.NeurIPS 2024 · 33 citations
