Feature-based Learning for Diverse and Privacy-Preserving Counterfactual Explanations
Vy Vo, Trung Le, Van Nguyen, He Zhao, Edwin V. Bonilla, Gholamreza Haffari, Dinh Q. Phung
Abstract
Interpretable machine learning seeks to understand the reasoning process of complex black-box systems that are long notorious for lack of explainability. One flourishing approach is through counterfactual explanations, which provide suggestions on what a user can do to alter an outcome. Not only must a counterfactual example counter the original prediction from the black-box classifier but it should also satisfy various constraints for practical applications. Diversity is one of the critical constraints that however remains less discussed. While diverse counterfactuals are ideal, it is computationally challenging to simultaneously address some other constraints. Furthermore, there is a growing privacy concern over the released counterfactual data. To this end, we propose a feature-based learning framework that effectively handles the counterfactual constraints and contributes itself to the limited pool of private explanation models. We demonstrate the flexibility and effectiveness of our method in generating diverse counterfactuals of actionability and plausibility. Our counterfactual engine is more efficient than counterparts of the same capacity while yielding the lowest re-identification risks. CCS CONCEPTS • Computing methodologies → Machine learning; • Security and privacy → Privacy protections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc287065-acb9-4400-b893-9f7e0b58bb51Cited by top-tier papers2
- From Counterfactuals to Trees: Competitive Analysis of Model Extraction AttacksAwa Khouna, Julien Ferry, Thibaut VidalNeurIPS 2025 · 3 citations
- AI2TALE: An Innovative Information Theory-based Approach for Learning to Localize Phishing AttacksVan Nguyen, Tingmin Wu, Xingliang Yuan, Marthie Grobler et al.ICLR 2025
Builds on6
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Numerical Composition of Differential PrivacySivakanth Gopi, Yin Tat Lee, Lukas WutschitzNeurIPS 2021 · 259 citations
- FOCUS: Flexible Optimizable Counterfactual Explanations for Tree EnsemblesAna Lucic, Harrie Oosterhuis, Hinda Haned, Maarten de RijkeAAAI 2022 · 87 citations
- Amortized Generation of Sequential Algorithmic Recourses for Black-Box ModelsSahil Verma, Keegan Hines, John P. DickersonAAAI 2022 · 28 citations
- Counterfactual Plans under Distributional AmbiguityNgoc Bui, Duy Nguyen, Viet Anh NguyenICLR 2022 · 26 citations
Related papers
- DiCoFlex: Model-Agnostic Diverse Counterfactuals with Flexible ControlOleksii Furman, Ulvi Movsum-zada, Patryk Marszalek, Maciej Zieba et al.NeurIPS 2025 · 3 citations
- Beyond Trivial Counterfactual Explanations with Diverse Valuable ExplanationsPau Rodríguez, Massimo Caccia, Alexandre Lacoste, Lee Zamparo et al.ICCV 2021 · 72 citations
- A General Search-Based Framework for Generating Textual Counterfactual ExplanationsDaniel Gilo, Shaul MarkovitchAAAI 2024 · 3 citations
- Model-Based Counterfactual Synthesizer for InterpretationFan Yang, Sahan Suresh Alva, Jiahao Chen, Xia HuKDD 2021 · 26 citations
- On Generating Plausible Counterfactual and Semi-Factual Explanations for Deep LearningEoin M. Kenny, Mark T. KeaneAAAI 2021 · 122 citations
