Interactive Inverse Reinforcement Learning for Cooperative Games
Thomas Kleine Büning, Anne-Marie George, Christos Dimitrakakis
摘要
We study the problem of designing autonomous agents that can learn to cooperate effectively with a potentially suboptimal partner while having no access to the joint reward function. This problem is modeled as a cooperative episodic two-agent Markov decision process. We assume control over only the first of the two agents in a Stackelberg formulation of the game, where the second agent is acting so as to maximise expected utility given the first agent's policy. How should the first agent act in order to learn the joint reward function as quickly as possible and so that the joint policy is as close to optimal as possible? We analyse how knowledge about the reward function can be gained in this interactive two-agent scenario. We show that when the learning agent's policies have a significant effect on the transition function, the reward function can be learned efficiently.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Environment Design for Inverse Reinforcement LearningThomas Kleine Buening, Victor Villin, Christos DimitrakakisICML 2024 · 被引用 5 次
- Belief-Driven Value Alignment for Human-Robot CollaborationSaisai Li, Bing Shi, Yiming Xia, Xiao SuAAAI 2026
它引用的顶会 Paper2
相关 Paper
- Learning to Steer Learners in GamesYizhou Zhang, Yian Ma, Eric MazumdarICML 2025
- Is Knowledge Power? On the (Im)possibility of Learning from Strategic InteractionsNivasini Ananthakrishnan, Nika Haghtalab, Chara Podimata, Kunhe YangNeurIPS 2024 · 被引用 9 次
- Cooperative Multi-player Bandit OptimizationIlai Bistritz, Nicholas BambosNeurIPS 2020 · 被引用 31 次
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等AAAI 2023 · 被引用 17 次
- Online Learning in Stackelberg Games with an Omniscient FollowerGeng Zhao, Banghua Zhu, Jiantao Jiao, Michael I. JordanICML 2023 · 被引用 23 次
