K-SHAP: Policy Clustering Algorithm for Anonymous Multi-Agent State-Action Pairs
Andrea Coletta, Svitlana Vyetrenko, Tucker Balch
Abstract
Learning agent behaviors from observational data has shown to improve our understanding of their decision-making processes, advancing our ability to explain their interactions with the environment and other agents. While multiple learning techniques have been proposed in the literature, there is one particular setting that has not been explored yet: multi agent systems where agent identities remain anonymous. For instance, in financial markets labeled data that identifies market participant strategies is typically proprietary, and only the anonymous state-action pairs that result from the interaction of multiple market participants are publicly available. As a result, sequences of agent actions are not observable, restricting the applicability of existing work. In this paper, we propose a Policy Clustering algorithm, called K-SHAP, that learns to group anonymous state-action pairs according to the agent policies. We frame the problem as an Imitation Learning (IL) task, and we learn a world-policy able to mimic all the agent behaviors upon different environmental states. We leverage the world-policy to explain each anonymous observation through an additive feature attribution method called SHAP (SHapley Additive exPlanations). Finally, by clustering the explanations we show that we are able to identify different agent policies and group observations accordingly. We evaluate our approach on simulated synthetic market data and a real-world financial dataset. We show that our proposal significantly and consistently outperforms the existing methods, identifying different agent strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81b2ee93-9529-49e2-bc72-e9ef5b520383Builds on5
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 72 citations
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 50 citations
- Top-Down Deep Clustering with Multi-Generator GANsDaniel P. M. de Mello, Renato M. Assunção, Fabricio MuraiAAAI 2022 · 22 citations
- Generative Attention Networks for Multi-Agent Behavioral ModelingMax Guangyu Li, Bo Jiang, Hao Zhu, Zhengping Che et al.AAAI 2020 · 21 citations
- TrafficSim: Learning To Simulate Realistic Multi-Agent BehaviorsSimon Suo, Sebastian Regalado, Sergio Casas, Raquel UrtasunCVPR 2021
Related papers
- SHERPA: Explainable Robust Algorithms for Privacy-Preserved Federated Learning in Future Networks to Defend Against Data Poisoning AttacksChamara Sandeepa, Bartlomiej Siniarski, Shen Wang, Madhusanka LiyanageS&P 2024 · 19 citations
- You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory PredictionOsama Makansi, Julius von Kügelgen, Francesco Locatello, Peter Vincent Gehler et al.ICLR 2022 · 35 citations
- Anonymous Bandits for Multi-User SystemsHossein Esfandiari, Vahab Mirrokni, Jon SchneiderNeurIPS 2022 · 2 citations
- Multi-Agent Interactions Modeling with Correlated PoliciesMinghuan Liu, Ming Zhou, Weinan Zhang, Yuzheng Zhuang et al.ICLR 2020 · 22 citations
- SHAP@k: Efficient and Probably Approximately Correct (PAC) Identification of Top-K FeaturesSanjay Kariyappa, Leonidas Tsepenekas, Freddy Lécué, Daniele MagazzeniAAAI 2024 · 11 citations
