UNIREX: A Unified Learning Framework for Language Model Rationale Extraction
Aaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan, Shaoliang Nie, Xiaochang Peng, Xiang Ren, Hamed Firooz
Abstract
An extractive rationale explains a language model’s (LM’s) prediction on a given task instance by highlighting the text inputs that most influenced the prediction. Ideally, rationale extraction should be faithful (reflective of LM’s actual behavior) and plausible (convincing to humans), without compromising the LM’s (i.e., task model’s) task performance. Although attribution algorithms and select-predict pipelines are commonly used in rationale extraction, they both rely on certain heuristics that hinder them from satisfying all three desiderata. In light of this, we propose UNIREX, a flexible learning framework which generalizes rationale extractor optimization as follows: (1) specify architecture for a learned rationale extractor; (2) select explainability objectives (i.e., faithfulness and plausibility criteria); and (3) jointly the train task model and rationale extractor on the task using selected objectives. UNIREX enables replacing prior works’ heuristic design choices with a generic learned rationale extractor in (1) and optimizing it for all three desiderata in (2)-(3). To facilitate comparison between methods w.r.t. multiple desiderata, we introduce the Normalized Relative Gain (NRG) metric. Across five English text classification datasets, our best UNIREX configuration outperforms the strongest baselines by an average of 32.9% NRG. Plus, we find that UNIREX-trained rationale extractors’ faithfulness can even generalize to unseen datasets and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02a17b1c-d76e-49f4-8ccc-5a46f410b6e7Cited by top-tier papers11
- D-Separation for Causal Self-ExplanationWei Liu, Jun Wang, Haozhao Wang, Ruixuan Li et al.NeurIPS 2023 · 29 citations
- Towards Trustworthy Explanation: On Causal RationalizationWenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai et al.ICML 2023 · 25 citations
- Tailoring Self-Rationalizers with Multi-Reward DistillationSahana Ramnath, Brihi Joshi, Skyler Hallinan, Ximing Lu et al.ICLR 2024 · 23 citations
- PINTO: Faithful Language Reasoning Using Prompt-Generated RationalesPeifeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen et al.ICLR 2023 · 21 citations
- Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-RationalizationWei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang et al.NeurIPS 2024 · 17 citations
Builds on9
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- Improving Deep Learning Interpretability by Saliency Guided TrainingAya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2021 · 121 citations
- Evaluating and Characterizing Human RationalesSamuel Carton, Anirudh Rathore, Chenhao TanEMNLP 2020 · 38 citations
- Discretized Integrated Gradients for Explaining Language ModelsSoumya Sanyal, Xiang RenEMNLP 2021 · 34 citations
- NILE : Natural Language Inference with Faithful Natural Language ExplanationsSawan Kumar, Partha P. TalukdarACL 2020 · 15 citations
Related papers
- Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsBodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, Julian J. McAuleyICML 2022 · 40 citations
- LIREx: Augmenting Language Inference with Relevant ExplanationsXinyan Zhao, V. G. Vinod VydiswaranAAAI 2021 · 41 citations
- Measuring Association Between Labels and Free-Text RationalesSarah Wiegreffe, Ana Marasovic, Noah A. SmithEMNLP 2021 · 12 citations
- QUASER: Question Answering with Scalable Extractive RationalizationAsish Ghoshal, Srinivasan Iyer, Bhargavi Paranjape, Kushal Lakhotia et al.SIGIR 2022 · 2 citations
- Making a (Counterfactual) Difference One Rationale at a TimeMitchell Plyler, Michael Green, Min ChiNeurIPS 2021 · 12 citations
