Consistent Counterfactuals for Deep Models
Emily Black, Zifan Wang, Matt Fredrikson
Abstract
Counterfactual examples are one of the most commonly-cited methods for explaining the predictions of machine learning models in key areas such as finance and medical diagnosis. Counterfactuals are often discussed under the assumption that the model on which they will be used is static, but in deployment models may be periodically retrained or fine-tuned. This paper studies the consistency of model prediction on counterfactual examples in deep networks under small changes to initial training conditions, such as weight initialization and leave-one-out variations in data, as often occurs during model deployment. We demonstrate experimentally that counterfactual examples for deep models are often inconsistent across such small changes, and that increasing the cost of the counterfactual, a stability-enhancing mitigation suggested by prior work in the context of simpler models, is not a reliable heuristic in deep networks. Rather, our analysis shows that a model's Lipschitz continuity around the counterfactual, along with confidence of its prediction, is key to its consistency across related models. To this end, we propose Stable Neighbor Search as a way to generate more consistent counterfactual explanations, and illustrate the effectiveness of this approach on several benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a301681f-bbc9-40b5-b882-330c3ce42c21Cited by top-tier papers11
- On the Adversarial Robustness of Causal Algorithmic RecourseRicardo Dominguez-Olmedo, Amir-Hossein Karimi, Bernhard SchölkopfICML 2022 · 80 citations
- Robust Counterfactual Explanations for Tree-Based EnsemblesSanghamitra Dutta, Jason Long, Saumitra Mishra, Cecilia Tilli et al.ICML 2022 · 73 citations
- Robust Counterfactual Explanations for Neural Networks With Probabilistic GuaranteesFaisal Hamman, Erfaun Noorani, Saumitra Mishra, Daniele Magazzeni et al.ICML 2023 · 54 citations
- Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope TheoryPasan Dissanayake, Sanghamitra DuttaNeurIPS 2024 · 17 citations
- Probabilistically Robust Recourse: Navigating the Trade-offs between Costs and Robustness in Algorithmic RecourseMartin Pawelczyk, Teresa Datta, Johannes van den Heuvel, Gjergji Kasneci et al.ICLR 2023 · 13 citations
Builds on4
- Algorithmic recourse under imperfect causal knowledge: a probabilistic approachAmir-Hossein Karimi, Bodo Julius von Kügelgen, Bernhard Schölkopf, Isabel ValeraNeurIPS 2020 · 224 citations
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 197 citations
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Fast Geometric Projections for Local Robustness CertificationAymeric Fromherz, Klas Leino, Matt Fredrikson, Bryan Parno et al.ICLR 2021 · 34 citations
Related papers
- Designing Counterfactual Generators using Deep Model InversionJayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Deepta Rajan, Jason Liang et al.NeurIPS 2021 · 25 citations
- Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model ChangeIgnacy Stepka, Jerzy Stefanowski, Mateusz LangoKDD 2025 · 1 citation
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow MatchingZhuo Cao, Xuan Zhao, Lena Krieger, Hanno Scharr et al.NeurIPS 2025 · 5 citations
- A General Search-Based Framework for Generating Textual Counterfactual ExplanationsDaniel Gilo, Shaul MarkovitchAAAI 2024 · 3 citations
- Is this the Right Neighborhood? Accurate and Query Efficient Model Agnostic ExplanationsAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Karthikeyan ShanmugamNeurIPS 2022 · 9 citations
