Entropy-Based Logic Explanations of Neural Networks
Pietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Pietro Lió, Marco Gori, Stefano Melacci
Abstract
Explainable artificial intelligence has rapidly emerged since lawmakers have started requiring interpretable models for safety-critical domains. Concept-based neural networks have arisen as explainable-by-design methods as they leverage human-understandable symbols (i.e. concepts) to predict class memberships. However, most of these approaches focus on the identification of the most relevant concepts but do not provide concise, formal explanations of how such concepts are leveraged by the classifier to make predictions. In this paper, we propose a novel end-to-end differentiable approach enabling the extraction of logic explanations from neural networks using the formalism of First-Order Logic. The method relies on an entropy-based criterion which automatically identifies the most relevant concepts. We consider four different case studies to demonstrate that: (i) this entropy-based criterion enables the distillation of concise logic explanations in safety-critical domains from clinical data to computer vision; (ii) the proposed approach outperforms state-of-the-art white-box models in terms of classification accuracy and matches black box performances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 88 citations
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga et al.ICML 2023 · 68 citations
- Global Concept-Based Interpretability for Graph Neural Networks via Neuron AnalysisHan Xuanyuan, Pietro Barbiero, Dobrik Georgiev, Lucie Charlotte Magister et al.AAAI 2023 · 62 citations
- Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic InterpretationsXinyue Xu, Yi Qin, Lu Mi, Hao Wang et al.ICLR 2024 · 32 citations
- Dividing and Conquering a BlackBox to a Mixture of Interpretable Models: Route, Interpret, RepeatShantanu Ghosh, Ke Yu, Forough Arabshahi, Kayhan BatmanghelichICML 2023 · 15 citations
Builds on3
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A Constraint-Based Approach to Learning and ExplanationGabriele Ciravegna, Francesco Giannini, Stefano Melacci, Marco Maggini et al.AAAI 2020 · 16 citations
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
Related papers
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng et al.CVPR 2021
- SIC: Similarity-Based Interpretable Image Classification with Neural NetworksTom Nuno Wolf, Emre Kavak, Fabian Bongratz, Christian WachingerICCV 2025 · 1 citation
- Formal Abductive Latent Explanations for Prototype-Based NetworksJules Soria, Zakaria Chihani, Julien Girard-Satabin, Alban Grastien et al.AAAI 2026 · 2 citations
- DeXAR: Deep Explainable Sensor-Based Activity Recognition in Smart-Home EnvironmentsLuca Arrotta, Gabriele Civitarese, Claudio BettiniUbiComp 2022 · 40 citations
- Learn to Explain Efficiently via Neural Logic Inductive LearningYuan Yang, Le SongICLR 2020 · 83 citations
