Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement Learning
Cheng Lu, Huayu Chen, Jianfei Chen, Hang Su, Chongxuan Li, Jun Zhu
Abstract
Guided sampling is a vital approach for applying diffusion models in real-world tasks that embeds human-defined guidance during the sampling procedure. This paper considers a general setting where the guidance is defined by an (unnormalized) energy function. The main challenge for this setting is that the intermediate guidance during the diffusion sampling procedure, which is jointly defined by the sampling distribution and the energy function, is unknown and is hard to estimate. To address this challenge, we propose an exact formulation of the intermediate guidance as well as a novel training objective named contrastive energy prediction (CEP) to learn the exact guidance. Our method is guaranteed to converge to the exact guidance under unlimited model capacity and data samples, while previous methods can not. We demonstrate the effectiveness of our method by applying it to offline reinforcement learning (RL). Extensive experiments on D4RL benchmarks demonstrate that our method outperforms existing state-of-the-art algorithms. We also provide some examples of applying CEP for image synthesis to demonstrate the scalability of CEP on high-dimensional data. Code is available at https://github.com/thu-ml/ CEP-energy-guided-diffusion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7abac11b-8338-4dde-ba2d-3f7513cc285eCited by top-tier papers93
- DiffusionNFT: Online Diffusion Reinforcement with Forward ProcessKaiwen Zheng, Huayu Chen, Haotian Ye, Haoxiang Wang et al.ICLR 2026 · 213 citations
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationShutong Ding, Ke Hu, Zhenhao Zhang, Kan Ren et al.NeurIPS 2024 · 132 citations
- TFG: Unified Training-Free Guidance for Diffusion ModelsHaotian Ye, Haowei Lin, Jiaqi Han, Minkai Xu et al.NeurIPS 2024 · 118 citations
- Noise Contrastive Alignment of Language Models with Explicit RewardsHuayu Chen, Guande He, Lifan Yuan, Ganqu Cui et al.NeurIPS 2024 · 103 citations
- Learning a Diffusion Model Policy from Rewards via Q-Score MatchingMichael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi MaICML 2024 · 90 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Analytic Energy-Guided Policy Optimization for Offline Reinforcement LearningJifeng Hu, Sili Huang, Zhejian Yang, Shengchao Hu et al.NeurIPS 2025 · 4 citations
- Energy-Weighted Flow Matching for Offline Reinforcement LearningShiyuan Zhang, Weitong Zhang, Quanquan GuICLR 2025
- Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement LearningXu-Hui Liu, Tian-Shuo Liu, Shengyi Jiang, Ruifeng Chen et al.ICML 2024 · 10 citations
- How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?Xiaoyuan Cheng, Wenxuan Yuan, Boyang Li, Yuanchao Xu et al.ICML 2026
- ContraDiff: Planning Towards High Return States via Contrastive LearningYixiang Shan, Zhengbang Zhu, Ting Long, Qifan Liang et al.ICLR 2025
