On Gradient-like Explanation under a Black-box Setting: When Black-box Explanations Become as Good as White-box
Yi Cai, Gerhard Wunder
Abstract
Attribution methods shed light on the explainability of data-driven approaches such as deep learning models by uncovering the most influential features in a to-be-explained decision. While determining feature attributions via gradients delivers promising results, the internal access required for acquiring gradients can be impractical under safety concerns, thus limiting the applicability of gradient-based approaches. In response to such limited flexibility, this paper presents (gradient-estimation-based explanation), an approach that produces gradient-like explanations through only query-level access. The proposed approach holds a set of fundamental properties for attribution methods, which are mathematically rigorously proved, ensuring the quality of its explanations. In addition to the theoretical analysis, with a focus on image data, the experimental results empirically demonstrate the superiority of the proposed method over state-of-the-art black-box methods and its competitive performance compared to methods with full access.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fd49839-e2d4-4484-9724-31cac08c0373Cited by top-tier papers3
- Tackling the XAI Disagreement Problem with Adaptive Feature GroupingGabriel Laberge, Ola AhmadICLR 2026
- Rethinking Explanation Evaluation Under the Retraining SchemeYi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard WunderAAAI 2026
- GEFA: A General Feature Attribution Framework Using Proxy Gradient EstimationYi Cai, Thibaud Ardoin, Gerhard WunderICML 2025
Related papers
- Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIWon Jun Kim, Hyungjin Chung, Jaemin Kim, Sangmin Lee et al.CVPR 2025
- Learning Deep Attribution Priors Based On Prior KnowledgeEthan Weinberger, Joseph D. Janizek, Su-In LeeNeurIPS 2020 · 27 citations
- Saliency strikes back: How filtering out high frequencies improves white-box explanationsSabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol et al.ICML 2024 · 4 citations
- MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang et al.AAAI 2024 · 16 citations
- Manifold Integrated Gradients: Riemannian Geometry for Feature AttributionEslam Zaher, Maciej Trzaskowski, Quan Nguyen, Fred RoostaICML 2024 · 13 citations
