SOLAR for Offline MARL: Plateau-Triggered Potential Shaping under World-Model Uncertainty
Jusheng Zhang, Yijia Fan, Ruiqi Chen, Jing Yang, Ziliang Chen, Yongsen Zheng, Yanxi Chen, Jian Wang, Kwok Yan Lam, Liang Lin, Keze Wang
摘要
Reward shaping can accelerate reinforcement learning, but in sparse-reward offline multi-agent RL it is often brittle: dense intrinsic rewards may alter the underlying Markov game, while world-model guidance can amplify model bias. We find that shaping becomes reliable when it is (i) activated only after statistically validated learning plateaus and (ii) constrained to potential-based shaping, which preserves the task optimum. Motivated by this, we propose SOLAR, a simulate--evaluate--shape framework. A learned world model enables low-cost rollouts to test plateaus; once a plateau is detected, we inject shaping in the form with adaptively updated potentials; and we attenuate shaping using uncertainty-aware throttling in unreliable regions. We provide theoretical analysis on policy invariance and on the deviation of plateau decisions under model error, and establish stability for the resulting two-timescale adaptation. Experiments on sparse-reward offline MARL benchmarks show consistent gains in stability and final performance across dataset qualities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningYiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng 等NeurIPS 2021 · 被引用 133 次
- MADiff: Offline Multi-agent Learning with Diffusion ModelsZhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang 等NeurIPS 2024 · 被引用 116 次
- Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor RectificationLing Pan, Longbo Huang, Tengyu Ma, Huazhe XuICML 2022 · 被引用 80 次
相关 Paper
- Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse RewardsRati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2022 · 被引用 122 次
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等AAAI 2023 · 被引用 17 次
- Bootstrapped Reward ShapingJacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. KulkarniAAAI 2025 · 被引用 7 次
- Reward Shaping Control Variates for Off-Policy Evaluation Under Sparse RewardsRitam Majumdar, Finale Doshi-Velez, Sonali ParbhooICML 2026 · 被引用 3 次
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima 等NeurIPS 2025 · 被引用 9 次
