Online Tuning for Offline Decentralized Multi-Agent Reinforcement Learning
Jiechuan Jiang, Zongqing Lu
Abstract
Offline reinforcement learning could learn effective policies from a fixed dataset, which is promising for real-world applications. However, in offline decentralized multi-agent reinforcement learning, due to the discrepancy between the behavior policy and learned policy, the transition dynamics in offline experiences do not accord with the transition dynamics in online execution, which creates severe errors in value estimates, leading to uncoordinated low-performing policies. One way to overcome this problem is to bridge offline training and online tuning. However, considering both deployment efficiency and sample efficiency, we could only collect very limited online experiences, making it insufficient to use merely online data for updating the agent policy. To utilize both offline and online experiences to tune the policies of agents, we introduce online transition correction (OTC) to implicitly correct the offline transition dynamics by modifying sampling probabilities. We design two types of distances, i.e., embedding-based and value-based distance, to measure the similarity between transitions, and further propose an adaptive rank-based prioritization to sample transitions according to the transition similarity. OTC is simple yet effective to increase data efficiency and improve agent policies in online tuning. Empirically, OTC outperforms baselines in a variety of tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bdcc4d0-9125-4b55-8a1b-e7fa2a0efebeCited by top-tier papers2
- Learning from Good Trajectories in Offline Multi-Agent Reinforcement LearningQi Tian, Kun Kuang, Furui Liu, Baoxiang WangAAAI 2023 · 14 citations
- Offline Opponent Modeling with Truncated Q-driven Instant Policy RefinementYuheng Jing, Kai Li, Bingyun Liu, Ziwen Zhang et al.ICML 2025
Builds on10
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind et al.ICML 2021 · 223 citations
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong et al.NeurIPS 2021 · 207 citations
Related papers
- Online Pre-Training for Offline-to-Online Reinforcement LearningYongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong et al.ICML 2025
- Actor-Critic Alignment for Offline-to-Online Reinforcement LearningZishun Yu, Xinhua ZhangICML 2023 · 50 citations
- Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangNeurIPS 2024 · 15 citations
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Adaptive Policy Learning for Offline-to-Online Reinforcement LearningHan Zheng, Xufang Luo, Pengfei Wei, Xuan Song et al.AAAI 2023 · 47 citations
