Learning to Collaborate with Unknown Agents in the Absence of Reward
Zuyuan Zhang, Hanhan Zhou, Mahdi Imani, Taeyoung Lee, Tian Lan
摘要
With the advancements of artificial intelligence (AI), emerging scenarios involving close collaboration between AI and other unknown agents are becoming increasingly common. This requires sometimes training AI agents to collaborate with unknown agents in the absence of a reward function -- which may be unavailable to the AI agents or even undefined by the unknown agents themselves -- thus posing news challenges to existing learning algorithms that often require knowing the shared reward. In this paper, we show that effective teaming with unknown agents can be achieved in the absence of a reward function, through actively modeling other unknown agents and reasoning about their latent rewards from available interaction/observation history. In particular, we propose a novel framework that leverages a kernel density Bayesian inverse learning method for active reward/goal inference and prove that multi-agent reinforcement learning guided by the inferred reward signals can converge to an optimal policy teaming with unknown agents. The result enables us to develop an adaptive policy update strategy, through the use of a family of pre-trained, goal-conditioned policies, further eliminating the need for online retraining. The proposed solution is evaluated using a wide range of diverse unknown agents of latent and even non-stationary reward. Our solution significantly increases the teaming performance between AI and unknown agents in the absence of reward.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree SearchSizhe Tang, Zuyuan Zhang, Mahdi Imani, Tian LanICML 2026 · 被引用 3 次
- MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active ElicitationZeyu Fang, Mahdi Imani, Tian LanICML 2026
它引用的顶会 Paper14
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
- Agent Modelling under Partial Observability for Deep Reinforcement LearningGeorgios Papoudakis, Filippos Christianos, Stefano V. AlbrechtNeurIPS 2021 · 被引用 110 次
- Maximum Entropy Population-Based Training for Zero-Shot Human-AI CoordinationRui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu 等AAAI 2023 · 被引用 94 次
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 被引用 90 次
相关 Paper
- Efficient Exploration of Reward Functions in Inverse Reinforcement Learning via Bayesian OptimizationSreejith Balakrishnan, Quoc Phong Nguyen, Bryan Kian Hsiang Low, Harold SohNeurIPS 2020 · 被引用 33 次
- Online Ad Hoc Teamwork under Partial ObservabilityPengjie Gu, Mengchen Zhao, Jianye Hao, Bo AnICLR 2022 · 被引用 35 次
- Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context VariablesYang Chen, Xiao Lin, Bo Yan, Libo Zhang 等AAAI 2024 · 被引用 8 次
- Scalable Bayesian Inverse Reinforcement LearningAlex James Chan, Mihaela van der SchaarICLR 2021 · 被引用 11 次
- Ad Hoc Teamwork via Offline Goal-Based Decision TransformersXinzhi Zhang, Hohei Chan, Deheng Ye, Yi Cai 等ICML 2025
