Enhancing the Rationale-Input Alignment for Self-explaining Rationalization
Wei Liu, Haozhao Wang, Jun Wang, Zhiying Deng, Yuankai Zhang, Cheng Wang, Ruixuan Li
摘要
Rationalization empowers deep learning models with self-explaining capabilities through a cooperative game, where a generator selects a semantically consistent subset of the input as a rationale, and a subsequent predictor makes predictions based on the selected rationale. In this paper, we discover that rationalization is prone to a problem named rationale shift, which arises from the algorithmic bias of the cooperative game. Rationale shift refers to a situation where the semantics of the selected rationale may deviate from the original input, but the predictor still produces accurate predictions based on the deviation, resulting in a compromised generator with misleading feedback. To address this issue, we first demonstrate the importance of the alignment between the rationale and the full input through both empirical observations and theoretical analysis. Subsequently, we introduce a novel approach called DAR (Discriminatively Aligned Rationalization), which utilizes an auxiliary module pretrained on the full input to discriminatively align the selected rationale and the original input. We theoretically illustrate how DAR accomplishes the desired alignment, thereby overcoming the rationale shift problem. The experiments on two widely used real-world benchmarks show that the proposed method significantly improves the explanation quality (measured by the overlap between the model-selected explanation and the human-annotated rationale) as compared to state-of-the-art techniques. Additionally, results on two synthetic settings further validate the effectiveness of DAR in addressing the rationale shift problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-RationalizationWei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang 等NeurIPS 2024 · 被引用 17 次
- GraphNarrator: Generating Textual Explanations for Graph Neural NetworksBo Pan, Zhen Xiong, Guanchen Wu, Zheng Zhang 等ACL 2025 · 被引用 7 次
- GNN Explanations that do not Explain and How to find ThemSteve Azzolin, Stefano Teso, Bruno Lepri, Andrea Passerini 等ICLR 2026 · 被引用 4 次
- Boosting Explainability through Selective Rationalization in Pre-trained Language ModelsLibing Yuan, Shuaibo Hu, Kui Yu, Le WuKDD 2025 · 被引用 1 次
- Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean DatasetsWei Liu, Zhongyu Niu, Lang Gao, Zhiying Deng 等ICML 2025
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Parameterized Explainer for Graph Neural NetworkDongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu 等NeurIPS 2020 · 被引用 888 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- Invariant RationalizationShiyu Chang, Yang Zhang, Mo Yu, Tommi S. JaakkolaICML 2020 · 被引用 232 次
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li 等AAAI 2020 · 被引用 113 次
相关 Paper
- Decoupled Rationalization with Asymmetric Learning Rates: A Flexible Lipschitz RestraintWei Liu, Jun Wang, Haozhao Wang, Ruixuan Li 等KDD 2023 · 被引用 4 次
- Learning Robust Rationales for Model Explainability: A Guidance-Based ApproachShuaibo Hu, Kui YuAAAI 2024 · 被引用 11 次
- Understanding Interlocking Dynamics of Cooperative RationalizationMo Yu, Yang Zhang, Shiyu Chang, Tommi S. JaakkolaNeurIPS 2021 · 被引用 52 次
- MGR: Multi-generator Based RationalizationWei Liu, Haozhao Wang, Jun Wang, Ruixuan Li 等ACL 2023 · 被引用 7 次
- Unsupervised Selective Rationalization with Noise InjectionAdam Storek, Melanie Subbiah, Kathleen R. McKeownACL 2023 · 被引用 4 次
