Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding
Xiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia, Gökcen Eraslan, Surag Nair, Tommaso Biancalani, Shuiwang Ji, Aviv Regev, Sergey Levine, Masatoshi Uehara
摘要
Diffusion models excel at capturing the natural design spaces of images, molecules, and biological sequences. However, for many applications, rather than merely generating designs that are natural, we aim to optimize downstream reward functions while preserving the naturalness of these design spaces. Existing methods for achieving this goal often require "differentiable" proxy models (e.g., classifier guidance) or computationally-expensive fine-tuning of diffusion models (e.g., classifier-free guidance, RL-based fine-tuning). Here, we propose a new method, Soft Value-based Decoding in Diffusion models (SVDD), to address these challenges. SVDD is an iterative sampling method that integrates soft value functions, which looks ahead to how intermediate noisy states lead to high rewards in the future, into the standard inference procedure of pre-trained diffusion models. Notably, SVDD avoids fine-tuning generative models and eliminates the need to construct differentiable models. This enables us to (1) directly use non-differentiable features/reward feedback, commonly used in many scientific domains, and (2) apply our method to recent discrete diffusion models in a principled way. Finally, we demonstrate the effectiveness of SVDD across several domains, including image generation, molecule generation (optimization of docking scores, QED, SA), and DNA/RNA generation (optimization of activity levels). The code is available at https://github.com/masa-ue/SVDD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper58
- Inference-time scaling of diffusion models through classical searchXiangcheng Zhang, Haowei Lin, Haotian Ye, James Y. Zou 等ICLR 2026 · 被引用 57 次
- Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsYinuo Ren, Haoxuan Chen, Yuchen Zhu, Wei Guo 等NeurIPS 2025 · 被引用 51 次
- Inference-Time Text-to-Video Alignment with Diffusion Latent Beam SearchYuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki FurutaNeurIPS 2025 · 被引用 50 次
- Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget ForcingJaihoon Kim, Taehoon Yoon, Jisung Hwang, Minhyuk SungNeurIPS 2025 · 被引用 43 次
- Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion modelsVineet Jain, Kusha Sareen, Mohammad Pedramfar, Siamak RavanbakhshNeurIPS 2025 · 被引用 41 次
它引用的顶会 Paper43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular DesignXingyu Su, Xiner Li, Masatoshi Uehara, Sunwoo Kim 等ICLR 2026 · 被引用 10 次
- Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein DesignChenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang 等ICLR 2025
- Value Matching: Scalable and Gradient-Free Reward-Guided Flow AdaptationCristian Perez Jensen, Luca Schaufelberger, Riccardo De Santi, Kjell Jorner 等ICLR 2026
- Directly Fine-Tuning Diffusion Models on Differentiable RewardsKevin Clark, Paul Vicol, Kevin Swersky, David J. FleetICLR 2024 · 被引用 377 次
- Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Ehsan Hajiramezanali, Gabriele Scalia 等NeurIPS 2024 · 被引用 31 次
