Reward-Directed Conditional Diffusion: Provable Distribution Estimation and Reward Improvement
Hui Yuan, Kaixuan Huang, Chengzhuo Ni, Minshuo Chen, Mengdi Wang
摘要
We explore the methodology and theory of reward-directed generation via conditional diffusion models. Directed generation aims to generate samples with desired properties as measured by a reward function, which has broad applications in generative AI, reinforcement learning, and computational biology. We consider the common learning scenario where the data set consists of unlabeled data along with a smaller set of data with noisy reward labels. Our approach leverages a learned reward function on the smaller data set as a pseudolabeler. From a theoretical standpoint, we show that this directed generator can effectively learn and sample from the reward-conditioned data distribution. Additionally, our model is capable of recovering the latent subspace representation of data. Moreover, we establish that the model generates a new population that moves closer to a user-specified target reward value, where the optimality gap aligns with the off-policy bandit regret in the feature subspace. The improvement in rewards obtained is influenced by the interplay between the strength of the reward signal, the distribution shift, and the cost of off-support extrapolation. We provide empirical results to validate our theory and highlight the relationship between the strength of extrapolation and the quality of generated samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Feedback Efficient Online Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali 等ICML 2024 · 被引用 47 次
- Unifying Generation and Prediction on Graphs with Latent Graph DiffusionCai Zhou, Xiyuan Wang, Muhan ZhangNeurIPS 2024 · 被引用 37 次
- Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Ehsan Hajiramezanali, Gabriele Scalia 等NeurIPS 2024 · 被引用 31 次
- Analysis of Learning a Flow-based Generative Model from Limited Sample ComplexityHugo Cui, Florent Krzakala, Eric Vanden-Eijnden, Lenka ZdeborováICLR 2024 · 被引用 31 次
- Constrained Diffusion Models via Dual TrainingShervin Khalafi, Dongsheng Ding, Alejandro RibeiroNeurIPS 2024 · 被引用 24 次
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement LearningZhendong Wang, Jonathan J. Hunt, Mingyuan ZhouICLR 2023 · 被引用 33 次
- Adding Conditional Control to Diffusion Models with Reinforcement LearningYulai Zhao, Masatoshi Uehara, Gabriele Scalia, Sun-Yuan Kung 等ICLR 2025 · 被引用 1 次
- Outsourced Diffusion Sampling: Efficient Posterior Inference in Latent Spaces of Generative ModelsSiddarth Venkatraman, Mohsin Hasan, Minsu Kim, Luca Scimeca 等ICML 2025
- Goal-directed Generation of Discrete Structures with Conditional Generative ModelsAmina Mollaysa, Brooks Paige, Alexandros KalousisNeurIPS 2020 · 被引用 12 次
- Training Diffusion Models Towards Diverse Image Generation with Reinforcement LearningZichen Miao, Jiang Wang, Ze Wang, Zhengyuan Yang 等CVPR 2024 · 被引用 12 次
