Learning to Collaborate with Unknown Agents in the Absence of Reward
Zuyuan Zhang, Hanhan Zhou, Mahdi Imani, Taeyoung Lee, Tian Lan
Abstract
With the advancements of artificial intelligence (AI), emerging scenarios involving close collaboration between AI and other unknown agents are becoming increasingly common. This requires sometimes training AI agents to collaborate with unknown agents in the absence of a reward function -- which may be unavailable to the AI agents or even undefined by the unknown agents themselves -- thus posing news challenges to existing learning algorithms that often require knowing the shared reward. In this paper, we show that effective teaming with unknown agents can be achieved in the absence of a reward function, through actively modeling other unknown agents and reasoning about their latent rewards from available interaction/observation history. In particular, we propose a novel framework that leverages a kernel density Bayesian inverse learning method for active reward/goal inference and prove that multi-agent reinforcement learning guided by the inferred reward signals can converge to an optimal policy teaming with unknown agents. The result enables us to develop an adaptive policy update strategy, through the use of a family of pre-trained, goal-conditioned policies, further eliminating the need for online retraining. The proposed solution is evaluated using a wide range of diverse unknown agents of latent and even non-stationary reward. Our solution significantly increases the teaming performance between AI and unknown agents in the absence of reward.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7d5be41-4889-4b76-a365-fccfe60972f4Cited by top-tier papers2
- NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree SearchSizhe Tang, Zuyuan Zhang, Mahdi Imani, Tian LanICML 2026 · 3 citations
- MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active ElicitationZeyu Fang, Mahdi Imani, Tian LanICML 2026
Builds on14
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand et al.ICML 2023 · 155 citations
- Agent Modelling under Partial Observability for Deep Reinforcement LearningGeorgios Papoudakis, Filippos Christianos, Stefano V. AlbrechtNeurIPS 2021 · 110 citations
- Maximum Entropy Population-Based Training for Zero-Shot Human-AI CoordinationRui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu et al.AAAI 2023 · 94 citations
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 90 citations
Related papers
- Efficient Exploration of Reward Functions in Inverse Reinforcement Learning via Bayesian OptimizationSreejith Balakrishnan, Quoc Phong Nguyen, Bryan Kian Hsiang Low, Harold SohNeurIPS 2020 · 33 citations
- Online Ad Hoc Teamwork under Partial ObservabilityPengjie Gu, Mengchen Zhao, Jianye Hao, Bo AnICLR 2022 · 35 citations
- Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context VariablesYang Chen, Xiao Lin, Bo Yan, Libo Zhang et al.AAAI 2024 · 8 citations
- Scalable Bayesian Inverse Reinforcement LearningAlex James Chan, Mihaela van der SchaarICLR 2021 · 11 citations
- Ad Hoc Teamwork via Offline Goal-Based Decision TransformersXinzhi Zhang, Hohei Chan, Deheng Ye, Yi Cai et al.ICML 2025
