Greedy when Sure and Conservative when Uncertain about the Opponents
Haobo Fu, Ye Tian, Hongxiang Yu, Weiming Liu, Shuang Wu, Jiechao Xiong, Ying Wen, Kai Li, Junliang Xing, Qiang Fu, Wei Yang
Abstract
We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response (conditional additionally on our opponent policy embedding) instead of a finite set of separate best responses against any opponent; 3) computes online a posterior of the current opponent policy embedding, without making the discrete and ineffective decision which type the current opponent belongs to; and 4) selects online between a real-time greedy policy and a fixed conservative policy via an adversarial bandit algorithm, gaining a theoretically better regret than adhering to either. Experimental studies on popular benchmarks demonstrate GSCU's superiority over the state-of-the-art methods. The code is available online at https: //github.com/YeTianJHU/GSCU .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bb0c045-e44e-4a12-92cc-d1f57e152ffaCited by top-tier papers14
- Policy Space Diversity for Non-Transitive GamesJian Yao, Weiming Liu, Haobo Fu, Yaodong Yang et al.NeurIPS 2023 · 28 citations
- A Robust and Opponent-Aware League Training Method for StarCraft IIRuozi Huang, Xipeng Wu, Hongsheng Yu, Zhong Fan et al.NeurIPS 2023 · 13 citations
- Stylized Offline Reinforcement Learning: Extracting Diverse High-Quality Behaviors from Heterogeneous DatasetsYihuan Mao, Chengjie Wu, Xi Chen, Hao Hu et al.ICLR 2024 · 9 citations
- Fast Peer Adaptation with Context-aware ExplorationLong Ma, Yuanfei Wang, Fangwei Zhong, Song-Chun Zhu et al.ICML 2024 · 9 citations
- Opponent Modeling with In-context SearchYuheng Jing, Bingyun Liu, Kai Li, Yifan Zang et al.NeurIPS 2024 · 8 citations
Builds on6
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- Agent Modelling under Partial Observability for Deep Reinforcement LearningGeorgios Papoudakis, Filippos Christianos, Stefano V. AlbrechtNeurIPS 2021 · 110 citations
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun et al.ICML 2021 · 66 citations
- Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information GameHaobo Fu, Weiming Liu, Shuang Wu, Yijia Wang et al.ICLR 2022 · 32 citations
- Fast Adaptation to New Environments via Policy-Dynamics Value FunctionsRoberta Raileanu, Maxwell Goldstein, Arthur Szlam, Rob FergusICML 2020 · 27 citations
Related papers
- Online Learning in Unknown Markov GamesYi Tian, Yuanhao Wang, Tiancheng Yu, Suvrit SraICML 2021 · 48 citations
- Latent Bandits RevisitedJoey Hong, Branislav Kveton, Manzil Zaheer, Yinlam Chow et al.NeurIPS 2020 · 55 citations
- Restless-UCB, an Efficient and Low-complexity Algorithm for Online Restless BanditsSiwei Wang, Longbo Huang, John C. S. LuiNeurIPS 2020 · 58 citations
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 241 citations
- Improved Algorithms for Conservative Exploration in BanditsEvrard Garcelon, Mohammad Ghavamzadeh, Alessandro Lazaric, Matteo PirottaAAAI 2020 · 24 citations
