Rethinking Optimal Transport in Offline Reinforcement Learning
Arip Asadulaev, Rostislav Korst, Aleksandr Korotin, Vage Egiazarian, Andrey Filchenkov, Evgeny Burnaev
摘要
We propose a novel algorithm for offline reinforcement learning using optimal transport. Typically, in offline reinforcement learning, the data is provided by various experts and some of them can be sub-optimal. To extract an efficient policy, it is necessary to stitch the best behaviors from the dataset. To address this problem, we rethink offline reinforcement learning as an optimal transportation problem. And based on this, we present an algorithm that aims to find a policy that maps states to a partial distribution of the best expert actions for each given state. We evaluate the performance of our algorithm on continuous control problems from the D4RL suite and demonstrate improvements over existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Reinforcement Learning via Value Gradient FlowHaoran Xu, Kaiwen Hu, Somayeh Sojoudi, Amy ZhangICLR 2026 · 被引用 4 次
- HOTA: Hamiltonian framework for Optimal Transport AdvectionNazar Buzun, Daniil Shlenskii, Maksim Bobrin, Dmitry V. DylovICLR 2026 · 被引用 2 次
- Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal TransportMingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin WangICML 2025
- AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly DetectionJunru Zhang, Lang Feng, Haoran Shi, Xu Guo 等ICML 2026
- Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal TransportYang Xiao, Weiming Liu, Jun Dan, Tengyue Xu 等CVPR 2026
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 被引用 350 次
- Optimal transport mapping via input convex neural networksAshok Vardhan Makkuva, Amirhossein Taghvaei, Sewoong Oh, Jason D. LeeICML 2020 · 被引用 254 次
相关 Paper
- Optimal Transport for Offline Imitation LearningYicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette 等ICLR 2023 · 被引用 2 次
- DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory StitchingGuanghe Li, Yixiang Shan, Zhengbang Zhu, Ting Long 等ICML 2024 · 被引用 41 次
- Adaptive Policy Learning for Offline-to-Online Reinforcement LearningHan Zheng, Xufang Luo, Pengfei Wei, Xuan Song 等AAAI 2023 · 被引用 47 次
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 被引用 121 次
- Cross-Domain Offline Policy Adaptation with Optimal Transport and Dataset ConstraintJiafei Lyu, Mengbei Yan, Zhongjian Qiao, Runze Liu 等ICLR 2025
