Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope Theory
Pasan Dissanayake, Sanghamitra Dutta
摘要
Counterfactual explanations provide ways of achieving a favorable model outcome with minimum input perturbation. However, counterfactual explanations can also be leveraged to reconstruct the model by strategically training a surrogate model to give similar predictions as the original (target) model. In this work, we analyze how model reconstruction using counterfactuals can be improved by further leveraging the fact that the counterfactuals also lie quite close to the decision boundary. Our main contribution is to derive novel theoretical relationships between the error in model reconstruction and the number of counterfactual queries required using polytope theory. Our theoretical analysis leads us to propose a strategy for model reconstruction that we call Counterfactual Clamping Attack (CCA) which trains a surrogate model using a unique loss function that treats counterfactuals differently than ordinary instances. Our approach also alleviates the related problem of decision boundary shift that arises in existing model reconstruction approaches when counterfactuals are treated as ordinary instances. Experimental results demonstrate that our strategy improves fidelity between the target and surrogate model predictions on several datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- From Counterfactuals to Trees: Competitive Analysis of Model Extraction AttacksAwa Khouna, Julien Ferry, Thibaut VidalNeurIPS 2025 · 被引用 3 次
- Few-Shot Knowledge Distillation of LLMs With Counterfactual ExplanationsFaisal Hamman, Pasan Dissanayake, Yanjun Fu, Sanghamitra DuttaNeurIPS 2025 · 被引用 3 次
- RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited DataXuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun 等ICML 2026
它引用的顶会 Paper12
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- ActiveThief: Model Extraction Using Active Learning and Unannotated Public DataSoham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade 等AAAI 2020 · 被引用 164 次
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 被引用 145 次
- Certified Monotonic Neural NetworksXingchao Liu, Xing Han, Na Zhang, Qiang LiuNeurIPS 2020 · 被引用 116 次
- Exploiting Explanations for Model Inversion AttacksXuejun Zhao, Wencan Zhang, Xiaokui Xiao, Brian Y. LimICCV 2021 · 被引用 113 次
相关 Paper
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- Learning Models for Actionable RecourseAlexis Ross, Himabindu Lakkaraju, Osbert BastaniNeurIPS 2021 · 被引用 25 次
- Reconstructing Training Data with Informed AdversariesBorja Balle, Giovanni Cherubin, Jamie HayesS&P 2022 · 被引用 214 次
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- Faithful Model Explanations through Energy-Constrained Conformal CounterfactualsPatrick Altmeyer, Mojtaba Farmanbar, Arie van Deursen, Cynthia C. S. LiemAAAI 2024 · 被引用 7 次
