An Empirical Study on Explanations in Out-of-Domain Settings
George Chrysostomou, Nikolaos Aletras
摘要
Recent work in Natural Language Processing has focused on developing approaches that extract faithful explanations, either via identifying the most important tokens in the input (i.e. post-hoc explanations) or by designing inherently faithful models that first select the most important tokens and then use them to predict the correct label (i.e. select-then-predict models). Currently, these approaches are largely evaluated on in-domain settings. Yet, little is known about how post-hoc explanations and inherently faithful models perform in out-ofdomain settings. In this paper, we conduct an extensive empirical study that examines: (1) the out-of-domain faithfulness of post-hoc explanations, generated by five feature attribution methods; and (2) the out-of-domain performance of two inherently faithful models over six datasets. Contrary to our expectations, results show that in many cases out-of-domain post-hoc explanation faithfulness measured by sufficiency and comprehensiveness is higher compared to in-domain. We find this misleading and suggest using a random baseline as a yardstick for evaluating post-hoc explanation faithfulness. Our findings also show that selectthen predict models demonstrate comparable predictive performance in out-of-domain settings to full-text trained models. 1 1 Code is attached to the submission and will be publicly released. 2 We use these terms interchangeably throughout our work. reasoning behind a model's prediction (Jacovi and 040 Goldberg, 2020) 041 Two popular methods for extracting explanations 042 are through feature attribution approaches (i.e. post-043 hoc explanation methods) or via inherently faithful 044 classifiers (i.e. select-then-predict models). The 045 first computes the contribution of different parts 046 of the input with respect to a model's prediction 047
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationJirui Qi, Gabriele Sarti, Raquel Fernández, Arianna BisazzaEMNLP 2024 · 被引用 6 次
- Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting TranscriptsHaochen Tan, Han Wu, Wei Shao, Xinyun Zhang 等EMNLP 2023 · 被引用 3 次
- Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions ExplainabilityJoakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo 等ACL 2025
它引用的顶会 Paper4
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
- Evaluating and Characterizing Human RationalesSamuel Carton, Anirudh Rathore, Chenhao TanEMNLP 2020 · 被引用 38 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
相关 Paper
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- Faithfulness Measurable Masked Language ModelsAndreas Madsen, Siva Reddy, Sarath ChandarICML 2024 · 被引用 6 次
- Provably Better Explanations with Optimized Aggregation of Feature AttributionsThomas Decker, Ananta R. Bhattarai, Jindong Gu, Volker Tresp 等ICML 2024 · 被引用 7 次
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm 等EMNLP 2022 · 被引用 29 次
- Incorporating Attribution Importance for Improving Faithfulness MetricsZhixue Zhao, Nikolaos AletrasACL 2023 · 被引用 4 次
