Trustworthy Actionable Perturbations
Jesse Friedbaum, Sudarshan Adiga, Ravi Tandon
摘要
Counterfactuals, or modified inputs that lead to a different outcome, are an important tool for understanding the logic used by machine learning classifiers and how to change an undesirable classification. Even if a counterfactual changes a classifier's decision, however, it may not affect the true underlying class probabilities, i.e. the counterfactual may act like an adversarial attack and ``fool'' the classifier. We propose a new framework for creating modified inputs that change the true underlying probabilities in a beneficial way which we call Trustworthy Actionable Perturbations (TAP). This includes a novel verification procedure to ensure that TAP change the true class probabilities instead of acting adversarially. Our framework also includes new cost, reward, and goal definitions that are better suited to effectuating change in the real world. We present PAC-learnability results for our verification procedure and theoretically analyze our new method for measuring reward. We also develop a methodology for creating TAP and compare our results to those achieved by previous counterfactual methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- From Search to Sampling: Generative Models for Robust Algorithmic RecoursePrateek Garg, Lokesh Nagalapatti, Sunita SarawagiICLR 2025
- Algorithmic Recourse for Long-Term ImprovementKentaro Kanamori, Ken Kobayashi, Satoshi Hara, Takuya TakagiICML 2025
它引用的顶会 Paper5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Algorithmic recourse under imperfect causal knowledge: a probabilistic approachAmir-Hossein Karimi, Bodo Julius von Kügelgen, Bernhard Schölkopf, Isabel ValeraNeurIPS 2020 · 被引用 224 次
- ML-LOO: Detecting Adversarial Examples with Feature AttributionPuyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang 等AAAI 2020 · 被引用 117 次
- Synthesizing Action Sequences for Modifying Model DecisionsGoutham Ramakrishnan, Yun Chan Lee, Aws AlbarghouthiAAAI 2020 · 被引用 37 次
- Improvement-Focused Causal Recourse (ICR)Gunnar König, Timo Freiesleben, Moritz Grosse-WentrupAAAI 2023 · 被引用 21 次
相关 Paper
- Ordered Counterfactual Explanation by Mixed-Integer Linear OptimizationKentaro Kanamori, Takuya Takagi, Ken Kobayashi, Yuichi Ike 等AAAI 2021 · 被引用 135 次
- Text Counterfactuals via Latent Optimization and Shapley-Guided SearchXiaoli Z. Fern, Quintin PopeEMNLP 2021 · 被引用 14 次
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- Learning Models for Actionable RecourseAlexis Ross, Himabindu Lakkaraju, Osbert BastaniNeurIPS 2021 · 被引用 25 次
- Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model ChangeIgnacy Stepka, Jerzy Stefanowski, Mateusz LangoKDD 2025 · 被引用 1 次
