Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope Theory
Pasan Dissanayake, Sanghamitra Dutta
Abstract
Counterfactual explanations provide ways of achieving a favorable model outcome with minimum input perturbation. However, counterfactual explanations can also be leveraged to reconstruct the model by strategically training a surrogate model to give similar predictions as the original (target) model. In this work, we analyze how model reconstruction using counterfactuals can be improved by further leveraging the fact that the counterfactuals also lie quite close to the decision boundary. Our main contribution is to derive novel theoretical relationships between the error in model reconstruction and the number of counterfactual queries required using polytope theory. Our theoretical analysis leads us to propose a strategy for model reconstruction that we call Counterfactual Clamping Attack (CCA) which trains a surrogate model using a unique loss function that treats counterfactuals differently than ordinary instances. Our approach also alleviates the related problem of decision boundary shift that arises in existing model reconstruction approaches when counterfactuals are treated as ordinary instances. Experimental results demonstrate that our strategy improves fidelity between the target and surrogate model predictions on several datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d52b0c76-d296-43d5-bd7c-44bd3615ea5fCited by top-tier papers3
- From Counterfactuals to Trees: Competitive Analysis of Model Extraction AttacksAwa Khouna, Julien Ferry, Thibaut VidalNeurIPS 2025 · 3 citations
- Few-Shot Knowledge Distillation of LLMs With Counterfactual ExplanationsFaisal Hamman, Pasan Dissanayake, Yanjun Fu, Sanghamitra DuttaNeurIPS 2025 · 3 citations
- RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited DataXuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun et al.ICML 2026
Builds on12
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- ActiveThief: Model Extraction Using Active Learning and Unannotated Public DataSoham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade et al.AAAI 2020 · 164 citations
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 145 citations
- Certified Monotonic Neural NetworksXingchao Liu, Xing Han, Na Zhang, Qiang LiuNeurIPS 2020 · 116 citations
- Exploiting Explanations for Model Inversion AttacksXuejun Zhao, Wencan Zhang, Xiaokui Xiao, Brian Y. LimICCV 2021 · 113 citations
Related papers
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- Learning Models for Actionable RecourseAlexis Ross, Himabindu Lakkaraju, Osbert BastaniNeurIPS 2021 · 25 citations
- Reconstructing Training Data with Informed AdversariesBorja Balle, Giovanni Cherubin, Jamie HayesS&P 2022 · 214 citations
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- Faithful Model Explanations through Energy-Constrained Conformal CounterfactualsPatrick Altmeyer, Mojtaba Farmanbar, Arie van Deursen, Cynthia C. S. LiemAAAI 2024 · 7 citations
