On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
Mathieu Serrurier, Franck Mamalet, Thomas Fel, Louis Béthune, Thibaut Boissin
Abstract
Input gradients have a pivotal role in a variety of applications, including adversarial attack algorithms for evaluating model robustness, explainable AI techniques for generating Saliency Maps, and counterfactual explanations. However, Saliency Maps generated by traditional neural networks are often noisy and provide limited insights. In this paper, we demonstrate that, on the contrary, the Saliency Maps of 1-Lipschitz neural networks, learned with the dual loss of an optimal transportation problem, exhibit desirable XAI properties: They are highly concentrated on the essential parts of the image with low noise, significantly outperforming state-of-theart explanation approaches across various models and metrics. We also prove that these maps align unprecedentedly well with human explanations on ImageNet. To explain the particularly beneficial properties of the Saliency Map for such models, we prove this gradient encodes both the direction of the transportation plan and the direction towards the nearest adversarial attack. Following the gradient down to the decision boundary is no longer considered an adversarial attack, but rather a counterfactual explanation that explicitly transports the input from one class to another. Thus, Learning with such a loss jointly optimizes the classification objective and the alignment of the gradient, i.e. the Saliency Map, to the transportation plan direction. These networks were previously known to be certifiably robust by design, and we demonstrate that they scale well for large problems and models, and are tailored for explainability using a fast and straightforward method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Vision Transformer Finetuning Benefits from Non-Smooth ComponentsAmbroise Odonnat, Chapel Laetitia, Romain Tavenard, Ievgen RedkoICML 2026 · 1 citation
- Universal Neural Optimal TransportJonathan Geuter, Gregor Kornhardt, Ingimar Tomasson, Vaios LaschosICML 2025
- Efficient Robust Conformal Prediction via Lipschitz-Bounded NetworksThomas Massena, Léo Andéol, Thibaut Boissin, Franck Mamalet et al.ICML 2025
- An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN ArchitecturesThibaut Boissin, Franck Mamalet, Thomas Fel, Agustin Martin Picard et al.ICML 2025
Builds on18
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 147 citations
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 137 citations
Related papers
- Achieving Robustness in Classification Using Optimal Transport With Hinge RegularizationMathieu Serrurier, Franck Mamalet, Alberto González-Sanz, Thibaut Boissin et al.CVPR 2021
- Building Reliable Explanations of Unreliable Neural Networks: Locally Smoothing Perspective of Model InterpretationDohun Lim, Hyeonseok Lee, Sungchan KimCVPR 2021
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
