K-SHAP: Policy Clustering Algorithm for Anonymous Multi-Agent State-Action Pairs
Andrea Coletta, Svitlana Vyetrenko, Tucker Balch
摘要
Learning agent behaviors from observational data has shown to improve our understanding of their decision-making processes, advancing our ability to explain their interactions with the environment and other agents. While multiple learning techniques have been proposed in the literature, there is one particular setting that has not been explored yet: multi agent systems where agent identities remain anonymous. For instance, in financial markets labeled data that identifies market participant strategies is typically proprietary, and only the anonymous state-action pairs that result from the interaction of multiple market participants are publicly available. As a result, sequences of agent actions are not observable, restricting the applicability of existing work. In this paper, we propose a Policy Clustering algorithm, called K-SHAP, that learns to group anonymous state-action pairs according to the agent policies. We frame the problem as an Imitation Learning (IL) task, and we learn a world-policy able to mimic all the agent behaviors upon different environmental states. We leverage the world-policy to explain each anonymous observation through an additive feature attribution method called SHAP (SHapley Additive exPlanations). Finally, by clustering the explanations we show that we are able to identify different agent policies and group observations accordingly. We evaluate our approach on simulated synthetic market data and a real-world financial dataset. We show that our proposal significantly and consistently outperforms the existing methods, identifying different agent strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 被引用 72 次
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 被引用 50 次
- Top-Down Deep Clustering with Multi-Generator GANsDaniel P. M. de Mello, Renato M. Assunção, Fabricio MuraiAAAI 2022 · 被引用 22 次
- Generative Attention Networks for Multi-Agent Behavioral ModelingMax Guangyu Li, Bo Jiang, Hao Zhu, Zhengping Che 等AAAI 2020 · 被引用 21 次
- TrafficSim: Learning To Simulate Realistic Multi-Agent BehaviorsSimon Suo, Sebastian Regalado, Sergio Casas, Raquel UrtasunCVPR 2021
相关 Paper
- SHERPA: Explainable Robust Algorithms for Privacy-Preserved Federated Learning in Future Networks to Defend Against Data Poisoning AttacksChamara Sandeepa, Bartlomiej Siniarski, Shen Wang, Madhusanka LiyanageS&P 2024 · 被引用 19 次
- You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory PredictionOsama Makansi, Julius von Kügelgen, Francesco Locatello, Peter Vincent Gehler 等ICLR 2022 · 被引用 35 次
- Anonymous Bandits for Multi-User SystemsHossein Esfandiari, Vahab Mirrokni, Jon SchneiderNeurIPS 2022 · 被引用 2 次
- Multi-Agent Interactions Modeling with Correlated PoliciesMinghuan Liu, Ming Zhou, Weinan Zhang, Yuzheng Zhuang 等ICLR 2020 · 被引用 22 次
- SHAP@k: Efficient and Probably Approximately Correct (PAC) Identification of Top-K FeaturesSanjay Kariyappa, Leonidas Tsepenekas, Freddy Lécué, Daniele MagazzeniAAAI 2024 · 被引用 11 次
