A Framework to Learn with Interpretation
Jayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc
摘要
With increasingly widespread use of deep neural networks in critical decision-making applications, interpretability of these models is becoming imperative. We consider the problem of jointly learning a predictive model and its associated interpretation model. The task of the interpreter is to provide both local and global interpretability about the predictive model in terms of human-understandable high level attribute functions, without any loss of accuracy. This is achieved by a dedicated architecture and well chosen regularization penalties. We seek for a small-size dictionary of attribute functions that take as inputs the outputs of selected hidden layers and whose outputs feed a linear classifier. We impose a high level of conciseness by constraining the activation of a very few attributes for a given input with a real-entropy-based criterion while enforcing fidelity to both inputs and outputs of the predictive model. A major advantage of simultaneous learning is that the predictive neural network benefits from the interpretability constraint as well. We also develop a more detailed pipeline based on some common and novel simple tools to develop understanding about the learnt features. We show on two datasets, MNIST and QuickDraw, their relevance for both global and local interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelSrishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin 等NeurIPS 2022 · 被引用 55 次
- Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMFJayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc 等NeurIPS 2022 · 被引用 32 次
- Time2Feat: Learning Interpretable Representations for Multivariate Time Series ClusteringAngela Bonifati, Francesco Del Buono, Francesco Guerra, Donato TianoVLDB 2023 · 被引用 27 次
- "Why Not Other Classes?": Towards Class-Contrastive Back-Propagation ExplanationsYipei Wang, Xiaoqian WangNeurIPS 2022 · 被引用 17 次
- TrajPAC: Towards Robustness Verification of Pedestrian Trajectory Prediction ModelsLiang Zhang, Nathaniel Xu, Pengfei Yang, Gaojie Jin 等ICCV 2023 · 被引用 13 次
它引用的顶会 Paper5
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang 等NeurIPS 2021 · 被引用 663 次
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- Learning outside the Black-Box: The pursuit of interpretable modelsJonathan Crabbé, Yao Zhang, William R. Zame, Mihaela van der SchaarNeurIPS 2020 · 被引用 30 次
- Learning to Faithfully Rationalize by ConstructionSarthak Jain, Sarah Wiegreffe, Yuval Pinter, Byron C. WallaceACL 2020
相关 Paper
- QPM: Discrete Optimization for Globally Interpretable Image ClassificationThomas Norrenbrock, Timo Kaiser, Sovan Biswas, Ramesh Manuvinakurike 等ICLR 2025
- Regional Tree Regularization for Interpretability in Deep Neural NetworksMike Wu, Sonali Parbhoo, Michael C. Hughes, Ryan Kindle 等AAAI 2020 · 被引用 42 次
- SOInter: A Novel Deep Energy-Based Interpretation Method for Explaining Structured Output ModelsSeyyede Fatemeh Seyyedsalehi, Mahdieh Soleymani Baghshah, Hamid R. RabieeICLR 2024
- Learning Global Transparent Models consistent with Local Contrastive ExplanationsTejaswini Pedapati, Avinash Balakrishnan, Karthikeyan Shanmugam, Amit DhurandharNeurIPS 2020 · 被引用 35 次
- Explainable Neural Networks with Guarantee: A Sparse Estimation ApproachAntoine Ledent, Peng LiuAAAI 2025 · 被引用 1 次
