A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
Anjie Liu, Jianhong Wang, Samuel Kaski, Jun Wang, Mengyue Yang
摘要
Steering cooperative multi-agent reinforcement learning (MARL) towards desired outcomes is challenging, particularly when the global guidance from a human on the whole multi-agent system is impractical in a large-scale MARL. On the other hand, designing external mechanisms (e.g., intrinsic rewards and human feedback) to coordinate agents mostly relies on empirical studies, lacking a easy-to-use research tool. In this work, we employ multi-agent influence diagrams (MAIDs) as a graphical framework to address the above issues. First, we introduce the concept of MARL interaction paradigms (orthogonal to MARL learning paradigms), using MAIDs to analyze and visualize both unguided self-organization and global guidance mechanisms in MARL. Then, we design a new MARL interaction paradigm, referred to as the targeted intervention paradigm that is applied to only a single targeted agent, so the problem of global guidance can be mitigated. In implementation, we introduce a causal inference technique, referred to as Pre-Strategy Intervention (PSI), to realize the targeted intervention paradigm. Since MAIDs can be regarded as a special class of causal diagrams, a composite desired outcome that integrates the primary task goal and an additional desired outcome can be achieved by maximizing the corresponding causal effect through the PSI. Moreover, the bundled relevance graph analysis of MAIDs provides a tool to identify whether an MARL learning paradigm is workable under the design of an MARL interaction paradigm. In experiments, we demonstrate the effectiveness of our proposed targeted intervention, and verify the result of relevance graph analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- ROMA: Multi-Agent Reinforcement Learning with Emergent RolesTonghan Wang, Heng Dong, Victor R. Lesser, Chongjie ZhangICML 2020 · 被引用 286 次
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song 等NeurIPS 2021 · 被引用 216 次
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 被引用 159 次
相关 Paper
- Situation-Dependent Causal Influence-Based Cooperative Multi-Agent Reinforcement LearningXiao Du, Yutong Ye, Pengyu Zhang, Yaning Yang 等AAAI 2024 · 被引用 19 次
- TMAE: Learning Targeted Multi-Agent Exploration via Causal InferenceChuxiong Sun, Dunqi Yao, Rui Wang, Wenwen Qiang 等AAAI 2026
- Multi-Agent Incentive Communication via Decentralized Teammate ModelingLei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang 等AAAI 2022 · 被引用 104 次
- MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG DiscoveryDong Li, Zhengzhang Chen, Xujiang Zhao, Linlin Yu 等AAAI 2026
- Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Xingchen Li 等NeurIPS 2023 · 被引用 7 次
