Unsupervised Selective Rationalization with Noise Injection
Adam Storek, Melanie Subbiah, Kathleen R. McKeown
Abstract
A major issue with using deep learning models in sensitive applications is that they provide no explanation for their output. To address this problem, unsupervised selective rationalization produces rationales alongside predictions by chaining two jointly-trained components, a rationale generator and a predictor. Although this architecture guarantees that the prediction relies solely on the rationale, it does not ensure that the rationale contains a plausible explanation for the prediction. We introduce a novel training technique that effectively limits generation of implausible rationales by injecting noise between the generator and the predictor. Furthermore, we propose a new benchmark for evaluating unsupervised selective rationalization models using movie reviews from existing datasets. We achieve sizeable improvements in rationale plausibility and task accuracy over the state-of-the-art across a variety of tasks, including our new benchmark, while maintaining or improving model faithfulness. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34f7a512-ace0-4a1d-8579-2236e56143e7Cited by top-tier papers4
- Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-RationalizationWei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang et al.NeurIPS 2024 · 17 citations
- Rationalizing Transformer Predictions via End-To-End Differentiable Self-TrainingMarc Felix Brinner, Sina ZarrießEMNLP 2024
- Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean DatasetsWei Liu, Zhongyu Niu, Lang Gao, Zhiying Deng et al.ICML 2025
- Breaking Free from MMI: A New Frontier in Rationalization by Probing Input UtilizationWei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang et al.ICLR 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Invariant RationalizationShiyu Chang, Yang Zhang, Mo Yu, Tommi S. JaakkolaICML 2020 · 232 citations
- Understanding Interlocking Dynamics of Cooperative RationalizationMo Yu, Yang Zhang, Shiyu Chang, Tommi S. JaakkolaNeurIPS 2021 · 52 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- Enhancing the Rationale-Input Alignment for Self-explaining RationalizationWei Liu, Haozhao Wang, Jun Wang, Zhiying Deng et al.ICDE 2024 · 6 citations
- Learning Robust Rationales for Model Explainability: A Guidance-Based ApproachShuaibo Hu, Kui YuAAAI 2024 · 11 citations
- Interlocking-free Selective Rationalization Through Genetic-based LearningFederico Ruggeri, Gaetano SignorelliACL 2025 · 1 citation
- Learning from the Best: Rationalizing Predictions by Adversarial Information CalibrationLei Sha, Oana-Maria Camburu, Thomas LukasiewiczAAAI 2021 · 40 citations
- Making a (Counterfactual) Difference One Rationale at a TimeMitchell Plyler, Michael Green, Min ChiNeurIPS 2021 · 12 citations
