Self-explaining deep models with logic rule reasoning
Seungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi, Xing Xie, Meeyoung Cha
Abstract
We present SELOR, a framework for integrating self-explaining capabilities into a given deep model to achieve both high prediction performance and human precision. By "human precision", we refer to the degree to which humans agree with the reasons models provide for their predictions. Human precision affects user trust and allows users to collaborate closely with the model. We demonstrate that logic rule explanations naturally satisfy human precision with the expressive power required for good predictive performance. We then illustrate how to enable a deep model to predict and explain with logic rules. Our method does not require predefined logic rule sets or human annotations and can be learned efficiently and easily with widely-used deep learning modules in a differentiable way. Extensive experiments show that our method gives explanations closer to human decision logic than other methods while maintaining the performance of deep learning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bc5b6d7-4def-4fc8-ada7-b9e7341f6ab3Cited by top-tier papers5
- Uncovering Safety Risks of Large Language Models through Concept Activation VectorZhihao Xu, Ruixuan Huang, Changyu Chen, Xiting WangNeurIPS 2024 · 83 citations
- A Causality Inspired Framework for Model InterpretationChenwang Wu, Xiting Wang, Defu Lian, Xing Xie et al.KDD 2023 · 22 citations
- Evaluating Readability and Faithfulness of Concept-based ExplanationsMeng Li, Haoran Jin, Ruixuan Huang, Zhihao Xu et al.EMNLP 2024 · 1 citation
- Shedding Light on Time Series Classification using Interpretability Gated NetworksYunshi Wen, Tengfei Ma, Ronny Luss, Debarun Bhattacharjya et al.ICLR 2025
- Explain Yourself, Briefly! Self-Explaining Neural Networks with Concise Sufficient ReasonsShahaf Bassan, Ron Eliav, Shlomit GurICLR 2025
Builds on8
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Human Factors in Model Interpretability: Industry Practices, Challenges, and NeedsSungsoo Ray Hong, Jessica Hullman, Enrico BertiniCSCW 2020 · 219 citations
- Multi-level Recommendation Reasoning over Knowledge Graphs with Reinforcement LearningXiting Wang, Kunpeng Liu, Dongjie Wang, Le Wu et al.WWW 2022 · 125 citations
- Leveraging Demonstrations for Reinforcement Recommendation Reasoning over Knowledge GraphsKangzhi Zhao, Xiting Wang, Yuren Zhang, Li Zhao et al.SIGIR 2020 · 114 citations
- FIND: Human-in-the-Loop Debugging Deep Text ClassifiersPiyawat Lertvittayakumjorn, Lucia Specia, Francesca ToniEMNLP 2020 · 33 citations
Related papers
- Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsBodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, Julian J. McAuleyICML 2022 · 40 citations
- Enhancing the Rationale-Input Alignment for Self-explaining RationalizationWei Liu, Haozhao Wang, Jun Wang, Zhiying Deng et al.ICDE 2024 · 6 citations
- Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable AttributesHaoyang Chen, Jingwen Bai, Fang Tian, Brian Y. LimCHI 2026 · 2 citations
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga et al.ICML 2023 · 68 citations
- A Framework for Learning Ante-hoc Explainable Models via ConceptsAnirban Sarkar, Deepak Vijaykeerthy, Anindya Sarkar, Vineeth N. BalasubramanianCVPR 2022 · 40 citations
