Understanding Interlocking Dynamics of Cooperative Rationalization
Mo Yu, Yang Zhang, Shiyu Chang, Tommi S. Jaakkola
Abstract
Selective rationalization explains the prediction of complex neural networks by finding a small subset of the input that is sufficient to predict the neural model output. The selection mechanism is commonly integrated into the model itself by specifying a two-component cascaded system consisting of a rationale generator, which makes a binary selection of the input features (which is the rationale), and a predictor, which predicts the output based only on the selected features. The components are trained jointly to optimize prediction performance. In this paper, we reveal a major problem with such cooperative rationalization paradigm -- model interlocking. Interlocking arises when the predictor overfits to the features selected by the generator thus reinforcing the generator's selection even if the selected rationales are sub-optimal. The fundamental cause of the interlocking problem is that the rationalization objective to be minimized is concave with respect to the generator's selection policy. We propose a new rationalization framework, called A2R, which introduces a third component into the architecture, a predictor driven by soft attention as opposed to selection. The generator now realizes both soft and hard attention over the features and these are fed into the two different predictors. While the generator still seeks to support the original predictor performance, it also minimizes a gap between the two predictors. As we will show theoretically, since the attention-based predictor exhibits a better convexity property, A2R can overcome the concavity barrier. Our experiments on two synthetic benchmarks and two real datasets demonstrate that A2R can significantly alleviate the interlock problem and find explanations that better align with human judgments. We release our code at https://github.com/Gorov/Understanding_Interlocking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f112b6a-3ba2-4708-ae37-9f2540dedeffCited by top-tier papers24
- Towards Interactivity and Interpretability: A Rationale-based Legal Judgment Prediction FrameworkYiquan Wu, Yifei Liu, Weiming Lu, Yating Zhang et al.EMNLP 2022 · 33 citations
- FR: Folded Rationalization with a Unified EncoderWei Liu, Haozhao Wang, Jun Wang, Ruixuan Li et al.NeurIPS 2022 · 33 citations
- D-Separation for Causal Self-ExplanationWei Liu, Jun Wang, Haozhao Wang, Ruixuan Li et al.NeurIPS 2023 · 29 citations
- Towards Trustworthy Explanation: On Causal RationalizationWenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai et al.ICML 2023 · 25 citations
- DARE: Disentanglement-Augmented Rationale ExtractionLinan Yue, Qi Liu, Yichao Du, Yanqing An et al.NeurIPS 2022 · 24 citations
Builds on5
- Invariant RationalizationShiyu Chang, Yang Zhang, Mo Yu, Tommi S. JaakkolaICML 2020 · 232 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Learning to Deceive with Attention-Based ExplanationsDanish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig et al.ACL 2020 · 17 citations
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra et al.ACL 2020 · 11 citations
- SPECTRA: Sparse Structured Text RationalizationNuno Miguel Guerreiro, André F. T. MartinsEMNLP 2021 · 1 citation
Related papers
- Interlocking-free Selective Rationalization Through Genetic-based LearningFederico Ruggeri, Gaetano SignorelliACL 2025 · 1 citation
- Learning Robust Rationales for Model Explainability: A Guidance-Based ApproachShuaibo Hu, Kui YuAAAI 2024 · 11 citations
- Interventional RationalizationLinan Yue, Qi Liu, Li Wang, Yanqing An et al.EMNLP 2023 · 8 citations
- Towards Faithful Explanations: Boosting Rationalization with Shortcuts DiscoveryLinan Yue, Qi Liu, Yichao Du, Li Wang et al.ICLR 2024 · 10 citations
- Enhancing the Rationale-Input Alignment for Self-explaining RationalizationWei Liu, Haozhao Wang, Jun Wang, Zhiying Deng et al.ICDE 2024 · 6 citations
