Feedback Efficient Online Fine-Tuning of Diffusion Models
Masatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali, Gabriele Scalia, Nathaniel Lee Diamant, Alex M. Tseng, Sergey Levine, Tommaso Biancalani
摘要
Diffusion models excel at modeling complex data distributions, including those of images, proteins, and small molecules. However, in many cases, our goal is to model parts of the distribution that maximize certain properties: for example, we may want to generate images with high aesthetic quality, or molecules with high bioactivity. It is natural to frame this as a reinforcement learning (RL) problem, in which the objective is to fine-tune a diffusion model to maximize a reward function that corresponds to some property. Even with access to online queries of the ground-truth reward function, efficiently discovering high-reward samples can be challenging: they might have a low probability in the initial distribution, and there might be many infeasible samples that do not even have a well-defined reward (e.g., unnatural images or physically impossible molecules). In this work, we propose a novel reinforcement learning procedure that efficiently explores on the manifold of feasible samples. We present a theoretical analysis providing a regret guarantee, as well as empirical validation across three domains: images, biological sequences, and molecules. The code is available at https://github.com/zhaoyl18/SEIKO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Aligning Target-Aware Molecule Diffusion Models with Exact Energy OptimizationSiyi Gu, Minkai Xu, Alexander S. Powers, Weili Nie 等NeurIPS 2024 · 被引用 35 次
- Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-TuningRiccardo De Santi, Marin Vlastelica, Ya-Ping Hsieh, Zebang Shen 等NeurIPS 2025 · 被引用 17 次
- Inference-Time Scaling of Discrete Diffusion Models via Importance Weighting and Optimal Proposal DesignZijing Ou, Chinmay Pani, Yingzhen LiICLR 2026 · 被引用 14 次
- Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based ModelsSangwoong Yoon, Himchan Hwang, Dohyun Kwon, Yung-Kyun Noh 等NeurIPS 2024 · 被引用 12 次
- Composition and Alignment of Diffusion Models using Constrained LearningShervin Khalafi, Ignacio Hounie, Dongsheng Ding, Alejandro RibeiroNeurIPS 2025 · 被引用 10 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- Uncertainty-Aware Multi-Objective Reinforcement Learning-Guided Diffusion Models for 3D De Novo Molecular DesignLianghong Chen, Dongkyu Eugene Kim, Mike Domaratzki, Pingzhao HuNeurIPS 2025 · 被引用 4 次
- Goal-directed Generation of Discrete Structures with Conditional Generative ModelsAmina Mollaysa, Brooks Paige, Alexandros KalousisNeurIPS 2020 · 被引用 12 次
- Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular DesignXingyu Su, Xiner Li, Masatoshi Uehara, Sunwoo Kim 等ICLR 2026 · 被引用 10 次
- Constrained Flow Optimization via Sequential Fine-Tuning for Molecular DesignSven Gutjahr, Riccardo De Santi, Luca Schaufelberger, Kjell Jorner 等ICML 2026 · 被引用 3 次
- Training Diffusion Models Towards Diverse Image Generation with Reinforcement LearningZichen Miao, Jiang Wang, Ze Wang, Zhengyuan Yang 等CVPR 2024 · 被引用 12 次
