A Framework to Learn with Interpretation
Jayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc
Abstract
With increasingly widespread use of deep neural networks in critical decision-making applications, interpretability of these models is becoming imperative. We consider the problem of jointly learning a predictive model and its associated interpretation model. The task of the interpreter is to provide both local and global interpretability about the predictive model in terms of human-understandable high level attribute functions, without any loss of accuracy. This is achieved by a dedicated architecture and well chosen regularization penalties. We seek for a small-size dictionary of attribute functions that take as inputs the outputs of selected hidden layers and whose outputs feed a linear classifier. We impose a high level of conciseness by constraining the activation of a very few attributes for a given input with a real-entropy-based criterion while enforcing fidelity to both inputs and outputs of the predictive model. A major advantage of simultaneous learning is that the predictive neural network benefits from the interpretability constraint as well. We also develop a more detailed pipeline based on some common and novel simple tools to develop understanding about the learnt features. We show on two datasets, MNIST and QuickDraw, their relevance for both global and local interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9aab19e4-2f72-42fa-94e5-5169c7fdbd58Cited by top-tier papers10
- ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelSrishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin et al.NeurIPS 2022 · 55 citations
- Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMFJayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc et al.NeurIPS 2022 · 32 citations
- Time2Feat: Learning Interpretable Representations for Multivariate Time Series ClusteringAngela Bonifati, Francesco Del Buono, Francesco Guerra, Donato TianoVLDB 2023 · 27 citations
- "Why Not Other Classes?": Towards Class-Contrastive Back-Propagation ExplanationsYipei Wang, Xiaoqian WangNeurIPS 2022 · 17 citations
- TrajPAC: Towards Robustness Verification of Pedestrian Trajectory Prediction ModelsLiang Zhang, Nathaniel Xu, Pengfei Yang, Gaojie Jin et al.ICCV 2023 · 13 citations
Builds on5
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang et al.NeurIPS 2021 · 663 citations
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Learning outside the Black-Box: The pursuit of interpretable modelsJonathan Crabbé, Yao Zhang, William R. Zame, Mihaela van der SchaarNeurIPS 2020 · 30 citations
- Learning to Faithfully Rationalize by ConstructionSarthak Jain, Sarah Wiegreffe, Yuval Pinter, Byron C. WallaceACL 2020
Related papers
- QPM: Discrete Optimization for Globally Interpretable Image ClassificationThomas Norrenbrock, Timo Kaiser, Sovan Biswas, Ramesh Manuvinakurike et al.ICLR 2025
- Regional Tree Regularization for Interpretability in Deep Neural NetworksMike Wu, Sonali Parbhoo, Michael C. Hughes, Ryan Kindle et al.AAAI 2020 · 42 citations
- SOInter: A Novel Deep Energy-Based Interpretation Method for Explaining Structured Output ModelsSeyyede Fatemeh Seyyedsalehi, Mahdieh Soleymani Baghshah, Hamid R. RabieeICLR 2024
- Learning Global Transparent Models consistent with Local Contrastive ExplanationsTejaswini Pedapati, Avinash Balakrishnan, Karthikeyan Shanmugam, Amit DhurandharNeurIPS 2020 · 35 citations
- Explainable Neural Networks with Guarantee: A Sparse Estimation ApproachAntoine Ledent, Peng LiuAAAI 2025 · 1 citation
