Pixel-level Certified Explanations via Randomized Smoothing
Alaa Anani, Tobias Lorenz, Mario Fritz, Bernt Schiele
Abstract
Post-hoc attribution methods aim to explain deep learning predictions by highlighting influential input pixels. However, these explanations are highly non-robust: small, imperceptible input perturbations can drastically alter the attribution map while maintaining the same prediction. This vulnerability undermines their trustworthiness and calls for rigorous robustness guarantees of pixellevel attribution scores. We introduce the first certification framework that guarantees pixel-level robustness for any black-box attribution method using randomized smoothing. By sparsifying and smoothing attribution maps, we reformulate the task as a segmentation problem and certify each pixel's importance against ℓ 2 -bounded perturbations. We further propose three evaluation metrics to assess certified robustness, localization, and faithfulness. An extensive evaluation of 12 attribution methods across 5 Ima-geNet models shows that our certified attributions are robust, interpretable, and faithful, enabling reliable use in downstream tasks. Our code is at https://github.com/AlaaAnani/ certified-attributions .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be7503ce-4929-48eb-9768-2920acf2deeaCited by top-tier papers5
- Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable GuaranteesItamar Hadad, Guy Katz, Shahaf BassanICLR 2026 · 10 citations
- Certified Circuits: Stability Guarantees for Mechanistic CircuitsAlaa Anani, Tobias Lorenz, Bernt Schiele, Mario Fritz et al.ICML 2026 · 3 citations
- Provably Explaining Neural Additive ModelsShahaf Bassan, Yizhak Yisrael Elboher, Tobias Ladner, Volkan Şahin et al.ICLR 2026 · 3 citations
- Unifying Formal Explanations: A Complexity-Theoretic PerspectiveShahaf Bassan, Xuanxiang Huang, Guy KatzICLR 2026 · 3 citations
- EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace MethodYanting Wang, Jinyuan JiaICLR 2026
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- On Locality of Local Explanation ModelsSahra Ghalebikesabi, Lucile Ter-Minassian, Karla DiazOrdaz, Chris C. HolmesNeurIPS 2021 · 52 citations
Related papers
- SAM: The Sensitivity of Attribution Methods to HyperparametersNaman Bansal, Chirag Agarwal, Anh NguyenCVPR 2020
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 2 citations
- Don't Lie to Me! Robust and Efficient Explainability with Verified Perturbation AnalysisThomas Fel, Melanie Ducoffe, David Vigouroux, Rémi Cadène et al.CVPR 2023
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 32 citations
- Manifold Integrated Gradients: Riemannian Geometry for Feature AttributionEslam Zaher, Maciej Trzaskowski, Quan Nguyen, Fred RoostaICML 2024 · 13 citations
