Learning to Scaffold: Optimizing Model Explanations for Teaching
Patrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins, Graham Neubig
摘要
Modern machine learning models are opaque, and as a result there is a burgeoning academic subfield on methods that explain these models' behavior. However, what is the precise goal of providing such explanations, and how can we demonstrate that explanations achieve this goal? Some research argues that explanations should help teach a student (either human or machine) to simulate the model being explained, and that the quality of explanations can be measured by the simulation accuracy of students on unexplained examples. In this work, leveraging metalearning techniques, we extend this idea to improve the quality of the explanations themselves, specifically by optimizing explanations such that student models more effectively learn to simulate the original model. We train models on three natural language processing and computer vision tasks, and find that students trained with explanations extracted with our framework are able to simulate the teacher significantly more effectively than ones produced with previous methods. Through human annotations and a user study, we further find that these learned explanations more closely align with how humans would explain the required decisions in these tasks. Our code is available at https://github.com/coderpat/learning-scaffold .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- D-Separation for Causal Self-ExplanationWei Liu, Jun Wang, Haozhao Wang, Ruixuan Li 等NeurIPS 2023 · 被引用 29 次
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 被引用 22 次
- Enhancing the Rationale-Input Alignment for Self-explaining RationalizationWei Liu, Haozhao Wang, Jun Wang, Zhiying Deng 等ICDE 2024 · 被引用 6 次
- Decoupled Rationalization with Asymmetric Learning Rates: A Flexible Lipschitz RestraintWei Liu, Jun Wang, Haozhao Wang, Ruixuan Li 等KDD 2023 · 被引用 4 次
- Breaking Free from MMI: A New Frontier in Rationalization by Probing Input UtilizationWei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang 等ICLR 2025
它引用的顶会 Paper16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
相关 Paper
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin 等ACL 2025 · 被引用 8 次
- Regularizing Black-box Models for Improved InterpretabilityGregory Plumb, Maruan Al-Shedivat, Ángel Alexander Cabrera, Adam Perer 等NeurIPS 2020 · 被引用 90 次
- Explainable Active Learning (XAL): Toward AI Explanations as Interfaces for Machine TeachersBhavya Ghai, Q. Vera Liao, Yunfeng Zhang, Rachel K. E. Bellamy 等CSCW 2020 · 被引用 107 次
- Self-explaining deep models with logic rule reasoningSeungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi 等NeurIPS 2022 · 被引用 27 次
- TVE: Learning Meta-attribution for Transferable Vision ExplainerGuanchu Wang, Yu-Neng Chuang, Fan Yang, Mengnan Du 等ICML 2024 · 被引用 1 次
