Self-explaining deep models with logic rule reasoning
Seungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi, Xing Xie, Meeyoung Cha
摘要
We present SELOR, a framework for integrating self-explaining capabilities into a given deep model to achieve both high prediction performance and human precision. By "human precision", we refer to the degree to which humans agree with the reasons models provide for their predictions. Human precision affects user trust and allows users to collaborate closely with the model. We demonstrate that logic rule explanations naturally satisfy human precision with the expressive power required for good predictive performance. We then illustrate how to enable a deep model to predict and explain with logic rules. Our method does not require predefined logic rule sets or human annotations and can be learned efficiently and easily with widely-used deep learning modules in a differentiable way. Extensive experiments show that our method gives explanations closer to human decision logic than other methods while maintaining the performance of deep learning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Uncovering Safety Risks of Large Language Models through Concept Activation VectorZhihao Xu, Ruixuan Huang, Changyu Chen, Xiting WangNeurIPS 2024 · 被引用 83 次
- A Causality Inspired Framework for Model InterpretationChenwang Wu, Xiting Wang, Defu Lian, Xing Xie 等KDD 2023 · 被引用 22 次
- Evaluating Readability and Faithfulness of Concept-based ExplanationsMeng Li, Haoran Jin, Ruixuan Huang, Zhihao Xu 等EMNLP 2024 · 被引用 1 次
- Shedding Light on Time Series Classification using Interpretability Gated NetworksYunshi Wen, Tengfei Ma, Ronny Luss, Debarun Bhattacharjya 等ICLR 2025
- Explain Yourself, Briefly! Self-Explaining Neural Networks with Concise Sufficient ReasonsShahaf Bassan, Ron Eliav, Shlomit GurICLR 2025
它引用的顶会 Paper8
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Human Factors in Model Interpretability: Industry Practices, Challenges, and NeedsSungsoo Ray Hong, Jessica Hullman, Enrico BertiniCSCW 2020 · 被引用 219 次
- Multi-level Recommendation Reasoning over Knowledge Graphs with Reinforcement LearningXiting Wang, Kunpeng Liu, Dongjie Wang, Le Wu 等WWW 2022 · 被引用 125 次
- Leveraging Demonstrations for Reinforcement Recommendation Reasoning over Knowledge GraphsKangzhi Zhao, Xiting Wang, Yuren Zhang, Li Zhao 等SIGIR 2020 · 被引用 114 次
- FIND: Human-in-the-Loop Debugging Deep Text ClassifiersPiyawat Lertvittayakumjorn, Lucia Specia, Francesca ToniEMNLP 2020 · 被引用 33 次
相关 Paper
- Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsBodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, Julian J. McAuleyICML 2022 · 被引用 40 次
- Enhancing the Rationale-Input Alignment for Self-explaining RationalizationWei Liu, Haozhao Wang, Jun Wang, Zhiying Deng 等ICDE 2024 · 被引用 6 次
- Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable AttributesHaoyang Chen, Jingwen Bai, Fang Tian, Brian Y. LimCHI 2026 · 被引用 2 次
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga 等ICML 2023 · 被引用 68 次
- A Framework for Learning Ante-hoc Explainable Models via ConceptsAnirban Sarkar, Deepak Vijaykeerthy, Anindya Sarkar, Vineeth N. BalasubramanianCVPR 2022 · 被引用 40 次
