Policy Improvement via Imitation of Multiple Oracles
Ching-An Cheng, Andrey Kolobov, Alekh Agarwal
摘要
Despite its promise, reinforcement learning's real-world adoption has been hampered by the need for costly exploration to learn a good policy. Imitation learning (IL) mitigates this shortcoming by using an oracle policy during training as a bootstrap to accelerate the learning process. However, in many practical situations, the learner has access to multiple suboptimal oracles, which may provide conflicting advice in a state. The existing IL literature provides a limited treatment of such scenarios. Whereas in the single-oracle case, the return of the oracle's policy provides an obvious benchmark for the learner to compete against, neither such a benchmark nor principled ways of outperforming it are known for the multi-oracle setting. In this paper, we propose the state-wise maximum of the oracle policies' values as a natural baseline to resolve conflicting advice from multiple oracles. Using a reduction of policy optimization to online learning, we introduce a novel IL algorithm MAMBA, which can provably learn a policy competitive with this benchmark. In particular, MAMBA optimizes policies by using a gradient estimator in the style of generalized advantage estimation (GAE). Our theoretical analysis shows that this design makes MAMBA robust and enables it to outperform the oracle policies by a larger margin than the IL state of the art, even in the single-oracle case. In an evaluation against standard policy gradient with GAE and AggreVaTe(D), we showcase MAMBA's ability to leverage demonstrations both from a single and from multiple weak oracles, and significantly speed up policy optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Adversarially Trained Actor Critic for Offline Reinforcement LearningChing-An Cheng, Tengyang Xie, Nan Jiang, Alekh AgarwalICML 2022 · 被引用 156 次
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 被引用 87 次
- Safe Reinforcement Learning Using Advantage-Based InterventionNolan Wagener, Byron Boots, Ching-An ChengICML 2021 · 被引用 66 次
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang 等NeurIPS 2023 · 被引用 41 次
相关 Paper
- Active Policy Improvement from Multiple Black-box OraclesXuefeng Liu, Takuma Yoneda, Chaoqi Wang, Matthew R. Walter 等ICML 2023 · 被引用 13 次
- Blending Imitation and Reinforcement Learning for Robust Policy ImprovementXuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter 等ICLR 2024 · 被引用 19 次
- Agnostic Interactive Imitation Learning: New Theory and Practical AlgorithmsYichen Li, Chicheng ZhangICML 2024
- IL-SOAR : Imitation Learning with Soft Optimistic Actor cRiticStefano Viel, Luca Viano, Volkan CevherICML 2025
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement LearningAlexander W. Goodall, Edwin Hamel-De le Court, Francesco BelardinelliAAAI 2026 · 被引用 1 次
