Planning with Quantized Opponent Models
Xiaopeng Yu, Kefan Su, Zongqing Lu
摘要
Planning under opponent uncertainty is a fundamental challenge in multi-agent environments, where an agent must act while inferring the hidden policies of its opponents. Existing type-based methods rely on manually defined behavior classes and struggle to scale, while model-free approaches are sample-inefficient and lack a principled way to incorporate uncertainty into planning. We propose Quantized Opponent Models (QOM), which learn a compact catalog of opponent types via a quantized autoencoder and maintain a Bayesian belief over these types online. This posterior supports both a belief-weighted meta-policy and a Monte-Carlo planning algorithm that directly integrates uncertainty, enabling real-time belief updates and focused exploration. Experiments show that QOM achieves superior performance with lower search cost, offering a tractable and effective solution for belief-aware planning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- Model-Based Opponent ModelingXiaopeng Yu, Jiechuan Jiang, Wanpeng Zhang, Haobin Jiang 等NeurIPS 2022 · 被引用 56 次
- OPtions as REsponses: Grounding behavioural hierarchies in multi-agent reinforcement learningAlexander Vezhnevets, Yuhuai Wu, Maria K. Eckstein, Rémi Leblond 等ICML 2020 · 被引用 44 次
相关 Paper
- Bayesian Optimized Monte Carlo PlanningJohn Mern, Anil Yildiz, Zachary Sunberg, Tapan Mukerji 等AAAI 2021 · 被引用 33 次
- A Bayesian Approach to Online PlanningNir Greshler, David Ben-Eli, Carmel Rabinovitz, Gabi Guetta 等ICML 2024 · 被引用 1 次
- Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and PlanningYizhe Huang, Anji Liu, Fanqi Kong, Yaodong Yang 等ICML 2024 · 被引用 5 次
- Scalable Decision-Making in Stochastic Environments through Learned Temporal AbstractionBaiting Luo, Ava Pettet, Aron Laszka, Abhishek Dubey 等ICLR 2025
- Improved Knowledge Modeling and Its Use for Signaling in Multi-Agent Planning with Partial ObservabilityShashank Shekhar, Ronen I. Brafman, Guy ShaniAAAI 2021 · 被引用 5 次
