Expected Returns and Policy Inconsistency-Aware Offline Federated Deep Reinforcement Learning
Meng XU, Zhongying Chen, Weiwei Fu, Yan Li, Shuguang Wang, Jianping Wang
摘要
Offline Federated Deep Reinforcement Learning (FDRL) methods aggregate multiple client-side offline Deep Reinforcement Learning (DRL) models, each trained locally, to facilitate knowledge sharing while preserving privacy. Existing offline FDRL methods assign client weights during global aggregation using either simple averaging or Q-values, but they neglect the combined consideration of Q-values and policy inconsistency, the latter of which reflects the distributional discrepancy between the learned policy and the policy from offline data. This causes clients with no significant advantages in one aspect but obvious disadvantages in the other to disproportionately affect the global model, thereby degrading its capabilities in that aspect. During local training, clients in existing methods are compelled to fully adopt the global model, which negatively impacts clients when the global model is weak. To this end, we propose a novel Federated Learning (FL) framework that can be seamlessly integrated into current offline FDRL approaches to improve their performance. Our method considers both policy inconsistency and Q-values to determine the weights of client models, with the latter adjusted by a scaling factor to avoid significant numerical discrepancies with the former. The aggregated global model is then distributed to clients to facilitate their learning from the global model. The impact of the global model on the local models is reduced when a client's model performance exceeds that of the global model, thereby mitigating the influence of a weaker global model. Experiments on the Datasets for Deep Data-Driven Reinforcement Learning (D4RL) demonstrate that our method improves seven state-of-the-art (SOTA) offline FDRL methods across several metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- On Convergence of FedProx: Local Dissimilarity Invariant Bounds, Non-smoothness and BeyondXiaotong Yuan, Ping LiNeurIPS 2022 · 被引用 141 次
- How to Leverage Unlabeled Data in Offline Reinforcement LearningTianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman 等ICML 2022 · 被引用 78 次
- Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced DatasetsZhang-Wei Hong, Aviral Kumar, Sathwik Karnik, Abhishek Bhandwaldar 等NeurIPS 2023 · 被引用 34 次
相关 Paper
- Federated Offline Policy Optimization with Dual RegularizationSheng Yue, Zerui Qin, Xingyuan Hua, Yongheng Deng 等INFOCOM 2024 · 被引用 2 次
- Federated Ensemble-Directed Offline Reinforcement LearningDesik Rengarajan, Nitin Ragothaman, Dileep Kalathil, Srinivas ShakkottaiNeurIPS 2024 · 被引用 11 次
- A Unified Self-Regulating Training Framework for Federated Deep Reinforcement LearningMeng Xu, Xinhong Chen, Zhongying Chen, Guanyi Zhao 等AAAI 2026
- Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage SufficesJiin Woo, Laixi Shi, Gauri Joshi, Yuejie ChiICML 2024 · 被引用 9 次
- On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared RepresentationsGuojun Xiong, Shufan Wang, Daniel Jiang, Jian LiICLR 2025
