Faithfulness Measurable Masked Language Models
Andreas Madsen, Siva Reddy, Sarath Chandar
摘要
A common approach to explaining NLP models is to use importance measures that express which tokens are important for a prediction. Unfortunately, such explanations are often wrong despite being persuasive. Therefore, it is essential to measure their faithfulness. One such metric is if tokens are truly important, then masking them should result in worse model performance. However, token masking introduces out-of-distribution issues, and existing solutions that address this are computationally expensive and employ proxy models. Furthermore, other metrics are very limited in scope. This work proposes an inherently faithfulness measurable model that addresses these challenges. This is achieved using a novel fine-tuning method that incorporates masking, such that masking tokens become in-distribution by design. This differs from existing approaches, which are completely model-agnostic but are inapplicable in practice. We demonstrate the generality of our approach by applying it to 16 different datasets and validate it using statistical in-distribution tests. The faithfulness is then measured with 9 different importance measures. Because masking is in-distribution, importance measures that themselves use masking become consistently more faithful. Additionally, because the model makes faithfulness cheap to measure, we can optimize explanations towards maximal faithfulness; thus, our model becomes indirectly inherently explainable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 被引用 121 次
- A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural NetworksMatan Haroush, Tzviel Frostig, Ruth Heller, Daniel SoudryICLR 2022 · 被引用 40 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Interpreting Language Models with Contrastive ExplanationsKayo Yin, Graham NeubigEMNLP 2022 · 被引用 32 次
相关 Paper
- Incorporating Attribution Importance for Improving Faithfulness MetricsZhixue Zhao, Nikolaos AletrasACL 2023 · 被引用 4 次
- NormXLogit: The Head-on-Top Never LiesSina Abbasi, Mohammad Reza Modarres, Mohammad Taher PilehvarEMNLP 2025
- An Empirical Study on Explanations in Out-of-Domain SettingsGeorge Chrysostomou, Nikolaos AletrasACL 2022
- Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language ModelsWei Jie Yeo, Ranjan Satapathy, Erik CambriaEMNLP 2025 · 被引用 2 次
- Explainable Token-level Noise Filtering for LLM Fine-tuning DatasetsYuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu 等ICLR 2026 · 被引用 1 次
