Trustworthy Actionable Perturbations
Jesse Friedbaum, Sudarshan Adiga, Ravi Tandon
Abstract
Counterfactuals, or modified inputs that lead to a different outcome, are an important tool for understanding the logic used by machine learning classifiers and how to change an undesirable classification. Even if a counterfactual changes a classifier's decision, however, it may not affect the true underlying class probabilities, i.e. the counterfactual may act like an adversarial attack and ``fool'' the classifier. We propose a new framework for creating modified inputs that change the true underlying probabilities in a beneficial way which we call Trustworthy Actionable Perturbations (TAP). This includes a novel verification procedure to ensure that TAP change the true class probabilities instead of acting adversarially. Our framework also includes new cost, reward, and goal definitions that are better suited to effectuating change in the real world. We present PAC-learnability results for our verification procedure and theoretically analyze our new method for measuring reward. We also develop a methodology for creating TAP and compare our results to those achieved by previous counterfactual methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25123c48-4c3e-42e6-81a7-d43f34ad44efCited by top-tier papers2
- From Search to Sampling: Generative Models for Robust Algorithmic RecoursePrateek Garg, Lokesh Nagalapatti, Sunita SarawagiICLR 2025
- Algorithmic Recourse for Long-Term ImprovementKentaro Kanamori, Ken Kobayashi, Satoshi Hara, Takuya TakagiICML 2025
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Algorithmic recourse under imperfect causal knowledge: a probabilistic approachAmir-Hossein Karimi, Bodo Julius von Kügelgen, Bernhard Schölkopf, Isabel ValeraNeurIPS 2020 · 224 citations
- ML-LOO: Detecting Adversarial Examples with Feature AttributionPuyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang et al.AAAI 2020 · 117 citations
- Synthesizing Action Sequences for Modifying Model DecisionsGoutham Ramakrishnan, Yun Chan Lee, Aws AlbarghouthiAAAI 2020 · 37 citations
- Improvement-Focused Causal Recourse (ICR)Gunnar König, Timo Freiesleben, Moritz Grosse-WentrupAAAI 2023 · 21 citations
Related papers
- Ordered Counterfactual Explanation by Mixed-Integer Linear OptimizationKentaro Kanamori, Takuya Takagi, Ken Kobayashi, Yuichi Ike et al.AAAI 2021 · 135 citations
- Text Counterfactuals via Latent Optimization and Shapley-Guided SearchXiaoli Z. Fern, Quintin PopeEMNLP 2021 · 14 citations
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- Learning Models for Actionable RecourseAlexis Ross, Himabindu Lakkaraju, Osbert BastaniNeurIPS 2021 · 25 citations
- Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model ChangeIgnacy Stepka, Jerzy Stefanowski, Mateusz LangoKDD 2025 · 1 citation
