Interpretable Neural Networks with Frank-Wolfe: Sparse Relevance Maps and Relevance Orderings
Jan MacDonald, Mathieu Besançon, Sebastian Pokutta
Abstract
We study the effects of constrained optimization formulations and Frank-Wolfe algorithms for obtaining interpretable neural network predictions. Reformulating the Rate-Distortion Explanations (RDE) method for relevance attribution as a constrained optimization problem provides precise control over the sparsity of relevance maps. This enables a novel multi-rate as well as a relevance-ordering variant of RDE that both empirically outperform standard RDE and other baseline methods in a well-established comparison test. We showcase several deterministic and stochastic variants of the Frank-Wolfe algorithm and their effectiveness for RDE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fedfc5bc-95e6-49bb-bbcc-a7a2453a5f81Cited by top-tier papers2
- Training Characteristic Functions with Reinforcement Learning: XAI-methods play Connect FourStephan Wäldchen, Sebastian Pokutta, Felix HuberICML 2022 · 9 citations
- Spatial-temporal Concept based Explanation of 3D ConvNetsYing Ji, Yu Wang, Jien KatoCVPR 2023
Builds on2
- Stochastic Frank-Wolfe for Constrained Finite-Sum MinimizationGeoffrey Négiar, Gideon Dresdner, Alicia Y. Tsai, Laurent El Ghaoui et al.ICML 2020 · 29 citations
- Simple steps are all you need: Frank-Wolfe and generalized self-concordant functionsAlejandro Carderera, Mathieu Besançon, Sebastian PokuttaNeurIPS 2021 · 22 citations
Related papers
- SPECTRA: Sparse Structured Text RationalizationNuno Miguel Guerreiro, André F. T. MartinsEMNLP 2021 · 1 citation
- Leveraging Sparse Linear Layers for Debuggable Deep NetworksEric Wong, Shibani Santurkar, Aleksander MadryICML 2021 · 101 citations
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 147 citations
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 22 citations
- Generating Attribution Maps with Disentangled Masked BackpropagationAdria Ruiz, Antonio Agudo, Francesc Moreno-NoguerICCV 2021 · 3 citations
