Training Characteristic Functions with Reinforcement Learning: XAI-methods play Connect Four
Stephan Wäldchen, Sebastian Pokutta, Felix Huber
摘要
One of the goals of Explainable AI (XAI) is to determine which input components were relevant for a classifier decision. This is commonly know as saliency attribution. Characteristic functions (from cooperative game theory) are able to evaluate partial inputs and form the basis for theoretically"fair"attribution methods like Shapley values. Given only a standard classifier function, it is unclear how partial input should be realised. Instead, most XAI-methods for black-box classifiers like neural networks consider counterfactual inputs that generally lie off-manifold. This makes them hard to evaluate and easy to manipulate. We propose a setup to directly train characteristic functions in the form of neural networks to play simple two-player games. We apply this to the game of Connect Four by randomly hiding colour information from our agents during training. This has three advantages for comparing XAI-methods: It alleviates the ambiguity about how to realise partial input, makes off-manifold evaluation unnecessary and allows us to compare the methods by letting them play against each other.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 被引用 799 次
- Counterfactual Explanations Can Be ManipulatedDylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer SinghNeurIPS 2021 · 被引用 182 次
- Shapley explainability on the data manifoldChristopher Frye, Damien de Mijolla, Tom Begley, Laurence Cowton 等ICLR 2021 · 被引用 125 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
- Interpretable Neural Networks with Frank-Wolfe: Sparse Relevance Maps and Relevance OrderingsJan MacDonald, Mathieu Besançon, Sebastian PokuttaICML 2022 · 被引用 13 次
相关 Paper
- Neural Payoff Machines: Predicting Fair and Stable Payoff Allocations Among Team MembersDaphne Cornelisse, Thomas Rood, Yoram Bachrach, Mateusz Malinowski 等NeurIPS 2022 · 被引用 10 次
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
- Towards Trustable SHAP ScoresOlivier Létoffé, Xuanxiang Huang, João Marques-SilvaAAAI 2025 · 被引用 23 次
- A Unifying Framework to the Analysis of Interaction Methods using Synergy FunctionsDaniel Lundström, Meisam RazaviyaynICML 2023 · 被引用 4 次
- Towards Attributions of Input Variables in a CoalitionXinhao Zheng, Huiqi Deng, Quanshi ZhangICML 2025
