VarFlow: Proper Scoring-Rule Diffusion Distillation via Energy Matching
Huiyang Shao, Xin Xia, Yuxi Ren, Xing Wang, Xuefeng Xiao
Abstract
Diffusion models achieve remarkable generative performance but are hampered by slow, iterative inference. Model distillation seeks to train a fast student generator. Variational Score Distillation (VSD) offers a principled KL-divergence minimization framework for this task. This method cleverly avoids computing the teacher model's Jacobian, but its student gradient relies on the score of the student's own noisy marginal distribution, ∇ xt log p ϕ,t (x t ). VSD thus requires approximations, such as training an auxiliary network to estimate this score. These approximations can introduce biases, cause training instability, or lead to an incomplete match of the target distribution, potentially focusing on conditional means rather than broader distributional features. We introduce VarFlow, a novel distillation method based on a framework we term Score-Rule Variational Distillation (SRVD) framework. VarFlow trains a one-step generator g ϕ (z) by directly minimizing an energy distance (derived from the strictly proper energy score) between the student's induced noisy data distribution p ϕ,t (x t ) and the teacher's target noisy distribution q t (x t ). This objective is estimated entirely using samples from these two distributions. Crucially, VarFlow bypasses the need to compute or approximate the intractable student score. By directly matching the full noisy marginal distributions, VarFlow aims for a more comprehensive and robust alignment between student and teacher, offering an efficient and theoretically grounded path to high-fidelity one-step generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d53840b-222a-49eb-a2e0-e983b3b86361Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Score-of-Mixture Training: One-Step Generative Model Training Made Simple via Score Estimation of Mixture DistributionsTejas Jayashankar, Jongha Jon Ryu, Gregory W. WornellICML 2025
- Mean Flow Distillation: Robust and Stable Distillation for Flow Matching ModelsAn Zhao, Shengyuan Zhang, Zhongjian Sun, Yixiang Zhou et al.ICML 2026 · 2 citations
- Noise Conditional Variational Score DistillationXinyu Peng, Ziyang Zheng, Yaoming Wang, Han Li et al.ICML 2025
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman et al.CVPR 2024 · 75 citations
- EM Distillation for One-step Diffusion ModelsSirui Xie, Zhisheng Xiao, Diederik P. Kingma, Tingbo Hou et al.NeurIPS 2024 · 69 citations
