Learn from A Rationalist: Distilling Intermediate Interpretable Rationales
Jiayi Dai, Randy Goebel
摘要
Because of the pervasive use of deep neural networks (DNNs), especially in high-stakes domains, the interpretability of DNNs has received increased attention. The general idea of rationale extraction (RE) is to provide an interpretable-by-design framework for DNNs via a select-predict architecture where two neural networks learn jointly to perform feature selection and prediction, respectively. Given only the remote supervision from the final task prediction, the process of learning to select subsets of features (or rationales ) requires searching in the space of all possible feature combinations, which is computationally challenging and even harder when the base neural networks are not sufficiently capable. To improve the predictive performance of RE models that are based on less capable or smaller neural networks (i.e., the students), we propose REKD ( R ationale E xtraction with K nowledge D istillation) where a student RE model learns from the rationales and predictions of a teacher (i.e., a rationalist ) in addition to the student's own RE optimization. This structural adjustment to RE aligns well with how humans could learn effectively from interpretable and verifiable knowledge. Because of the neural-model agnostic nature of the method, any black-box neural network could be integrated as a backbone model. To demonstrate the viability of REKD, we conduct experiments with multiple variants of BERT and vision transformer (ViT) models. Our experiments across language and vision classification datasets (i.e., IMDB movie reviews, CIFAR 10 and CIFAR 100) show that REKD significantly improves the predictive performance of the student RE models. The code is publicly available: https://github.com/JiayiDai/REKD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Understanding Interlocking Dynamics of Cooperative RationalizationMo Yu, Yang Zhang, Shiyu Chang, Tommi S. JaakkolaNeurIPS 2021 · 被引用 52 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
相关 Paper
- Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsBodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, Julian J. McAuleyICML 2022 · 被引用 40 次
- Learning from the Best: Rationalizing Predictions by Adversarial Information CalibrationLei Sha, Oana-Maria Camburu, Thomas LukasiewiczAAAI 2021 · 被引用 40 次
- Self-training with Few-shot RationalizationMeghana Moorthy Bhat, Alessandro Sordoni, Subhabrata MukherjeeEMNLP 2021
- Learning to Rationalize for Nonmonotonic Reasoning with Distant SupervisionFaeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin ChoiAAAI 2021 · 被引用 46 次
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng 等CVPR 2021
