VarFlow: Proper Scoring-Rule Diffusion Distillation via Energy Matching
Huiyang Shao, Xin Xia, Yuxi Ren, Xing Wang, Xuefeng Xiao
摘要
Diffusion models achieve remarkable generative performance but are hampered by slow, iterative inference. Model distillation seeks to train a fast student generator. Variational Score Distillation (VSD) offers a principled KL-divergence minimization framework for this task. This method cleverly avoids computing the teacher model's Jacobian, but its student gradient relies on the score of the student's own noisy marginal distribution, ∇ xt log p ϕ,t (x t ). VSD thus requires approximations, such as training an auxiliary network to estimate this score. These approximations can introduce biases, cause training instability, or lead to an incomplete match of the target distribution, potentially focusing on conditional means rather than broader distributional features. We introduce VarFlow, a novel distillation method based on a framework we term Score-Rule Variational Distillation (SRVD) framework. VarFlow trains a one-step generator g ϕ (z) by directly minimizing an energy distance (derived from the strictly proper energy score) between the student's induced noisy data distribution p ϕ,t (x t ) and the teacher's target noisy distribution q t (x t ). This objective is estimated entirely using samples from these two distributions. Crucially, VarFlow bypasses the need to compute or approximate the intractable student score. By directly matching the full noisy marginal distributions, VarFlow aims for a more comprehensive and robust alignment between student and teacher, offering an efficient and theoretically grounded path to high-fidelity one-step generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Score-of-Mixture Training: One-Step Generative Model Training Made Simple via Score Estimation of Mixture DistributionsTejas Jayashankar, Jongha Jon Ryu, Gregory W. WornellICML 2025
- Mean Flow Distillation: Robust and Stable Distillation for Flow Matching ModelsAn Zhao, Shengyuan Zhang, Zhongjian Sun, Yixiang Zhou 等ICML 2026 · 被引用 2 次
- Noise Conditional Variational Score DistillationXinyu Peng, Ziyang Zheng, Yaoming Wang, Han Li 等ICML 2025
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman 等CVPR 2024 · 被引用 75 次
- EM Distillation for One-step Diffusion ModelsSirui Xie, Zhisheng Xiao, Diederik P. Kingma, Tingbo Hou 等NeurIPS 2024 · 被引用 69 次
