Explanation-based Finetuning Makes Models More Robust to Spurious Cues
Josh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah, Qing Lyu, Marianna Apidianaki, Chris Callison-Burch
Abstract
Large Language Models (LLMs) are so powerful that they sometimes learn correlations between labels and features that are irrelevant to the task, leading to poor generalization on outof-distribution data. We propose explanationbased finetuning as a general approach to mitigate LLMs' reliance on spurious correlations. Unlike standard finetuning where the model only predicts the answer given the input, we finetune the model to additionally generate a free-text explanation supporting its answer. To evaluate our method, we finetune the model on artificially constructed training sets containing different types of spurious cues, and test it on a test set without these cues. Compared to standard finetuning, our method makes GPT-3 (davinci) remarkably more robust against spurious cues in terms of accuracy drop across four classification tasks: ComVE (+1.2), CREAK (+9.1), e-SNLI (+15.4), and SBIC (+6.5). The efficacy generalizes across multiple model families and scales, with greater gains for larger models. Finally, our method also works well with explanations generated by the model, implying its applicability to more datasets without human-written explanations. 1,2 * Equal contribution. 1 Warning: this paper contains examples that may be offensive or upsetting. 2 Our code is available at https://github.com/ taidnguyen/explanation-based_finetuning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Using Natural Language Explanations to Improve Robustness of In-context LearningXuanli He, Yuxiang Wu, Oana-Maria Camburu, Pasquale Minervini et al.ACL 2024 · 7 citations
- ICLEF: In-Context Learning with Expert Feedback for Explainable Style TransferArkadiy Saakyan, Smaranda MuresanACL 2024
- Improving Implicit Discourse Relation Recognition with Natural Language Explanations from LLMsHeng Wang, Changxing WuAAAI 2026
- RORA: Robust Free-Text Rationale EvaluationZhengping Jiang, Yining Lu, Hanjie Chen, Daniel Khashabi et al.ACL 2024
- Exploring Explanations Improves the Robustness of In-Context LearningUkyo Honda, Tatsushi OkaACL 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 136 citations
Related papers
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 272 citations
- FLamE: Few-shot Learning from Natural Language ExplanationsYangqiaoyu Zhou, Yiming Zhang, Chenhao TanACL 2023 · 8 citations
- ArGue: Attribute-Guided Prompt Tuning for Vision-Language ModelsXinyu Tian, Shu Zou, Zhaoyuan Yang, Jing ZhangCVPR 2024 · 29 citations
- CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language ModelsKairong Han, Wenshuo Zhao, Ziyu Zhao, Ye Jun Jian et al.EMNLP 2025 · 3 citations
- Multi-Level Explanations for Generative Language ModelsLucas Monteiro Paes, Dennis Wei, Hyo Jin Do, Hendrik Strobelt et al.ACL 2025 · 16 citations
