Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
Yuanzhi Zhu, Xi Wang, Stéphane Lathuilière, Vicky Kalogeiton
摘要
One-step generators distilled from Masked Diffusion Models (MDMs) compress multiple sampling steps into a single forward pass, enabling efficient text and image synthesis. However, they suffer two key limitations: they inherit modeling bias from the teacher, and their discrete token outputs block gradient flow, preventing post-distillation refinements such as adversarial training, reward-based fine-tuning, and Test-Time Embedding Optimization (TTEO). In this work, we introduce soft embeddings, a simple relaxation that replaces discrete tokens with the expected embeddings under the generator's output distribution. Soft embeddings preserve representation fidelity for one-step discrete generator while providing a fully differentiable continuous surrogate that is compatible with teacher backbones and tokenizer decoders while cause minimum bias. Integrating soft embeddings into the Di[M]O distillation framework (denoted Soft-Di[M]O) makes one-step generators end-to-end trainable and enables straightforward application of GAN-based refinement, differentiable reward fine-tuning, and TTEO. Empirically, across multiple MDM teachers (e.g., MaskBit , MaskGen ), Soft-Di[M]O achieves state-of-the-art one-step results: improved class-to-image performance, a one-step FID of 1.56 on ImageNet-256 with GAN-based refinement, along with higher than teacher GenEval and HPS scores on text-to-image with reward fine-tuning, and further gains from TTEO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations ModelingTianyu Xie, Shuchen Xue, Zijin Feng, Tianyang Hu 等ICLR 2026 · 被引用 14 次
- LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language ModelsChenglin Wang, Yucheng Zhou, Shuang Chen, Tao Wang 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper86
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Di[M]O: Distilling Masked Diffusion Models Into One-Step GeneratorYuanzhi Zhu, Xi Wang, Stéphane Lathuilière, Vicky KalogeitonICCV 2025
- EM Distillation for One-step Diffusion ModelsSirui Xie, Zhisheng Xiao, Diederik P. Kingma, Tingbo Hou 等NeurIPS 2024 · 被引用 69 次
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang 等NeurIPS 2024 · 被引用 728 次
- Revisiting Diffusion Models: From Generative Pre-training to One-Step GenerationBowen Zheng, Tianming YangICML 2025
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score DistillationEnshu Liu, Qian Chen, Xuefei Ning, Shengen Yan 等NeurIPS 2025 · 被引用 4 次
