Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
Mingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin Wang
Abstract
Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and longterm planning. However, they face challenges in robustness when encountering distribution shifts. This paper explores improving diffusion-based imitation learning models through online interactions with the environment. We propose OTPR (Optimal Transport-guided score-based diffusion Policy for Reinforcement learning fine-tuning), a novel method that integrates diffusion policies with RL using optimal transport theory. OTPR leverages the Q-function as a transport cost and views the policy as an optimal transport map, enabling efficient and stable fine-tuning. Moreover, we introduce masked optimal transport to guide state-action matching using expert keypoints and a compatibility-based resampling strategy to enhance training stability. Experiments on three simulation tasks demonstrate OTPR's superior performance and robustness compared to existing methods, especially in complex and sparsereward environments. In sum, OTPR provides an effective framework for combining IL and RL, achieving versatile and reliable policy learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b233132e-2f3e-40f0-9142-a953272c52aaCited by top-tier papers1
Ask how each one uses itBuilds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Rethinking Optimal Transport in Offline Reinforcement LearningArip Asadulaev, Rostislav Korst, Aleksandr Korotin, Vage Egiazarian et al.NeurIPS 2024 · 12 citations
- Optimal Transport for Offline Imitation LearningYicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette et al.ICLR 2023 · 2 citations
- Efficient and Uncertainty-Aware Diffusion Framework for Offline-to-Online Reinforcement LearningHa Manh Bui, Metod Jazbec, Eric Nalisnick, Anqi LiuICML 2026
- OMPO: A Unified Framework for RL under Policy and Dynamics ShiftsYu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang et al.ICML 2024 · 5 citations
- Efficient Online Reinforcement Learning for Diffusion PolicyHaitong Ma, Tianyi Chen, Kai Wang, Na Li et al.ICML 2025
