Using Natural Language Explanations to Improve Robustness of In-context Learning
Xuanli He, Yuxiang Wu, Oana-Maria Camburu, Pasquale Minervini, Pontus Stenetorp
Abstract
Recent studies demonstrated that large language models (LLMs) can excel in many tasks via in-context learning (ICL). However, recent works show that ICL-prompted models tend to produce inaccurate results when presented with adversarial inputs. In this work, we investigate whether augmenting ICL with natural language explanations (NLEs) improves the robustness of LLMs on adversarial datasets covering natural language inference and paraphrasing identification. We prompt LLMs with a small set of human-generated NLEs to produce further NLEs, yielding more accurate results than both a zero-shot-ICL setting and using only human-generated NLEs. Our results on five popular LLMs (GPT3.5-turbo, Llama2, Vicuna, Zephyr, and Mistral) show that our approach yields over 6% improvement over baseline approaches for eight adversarial datasets: HANS, ISCS, NaN, ST, PICD, PISP, ANLI, and PAWS. Furthermore, previous studies have demonstrated that prompt selection strategies significantly enhance ICL on in-distribution test sets. However, our findings reveal that these strategies do not match the efficacy of our approach for robustness evaluations, resulting in an accuracy drop of 8% compared to the proposed approach. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support SettingMaxime Kayser, Bayar Menzat, Cornelius Emde, Bogdan Bercean et al.EMNLP 2024 · 8 citations
- Faithful and Robust LLM-Driven Theorem Proving for NLI ExplanationsXin Quan, Marco Valentino, Louise A. Dennis, André FreitasACL 2025 · 8 citations
- PCoT: Persuasion-Augmented Chain of Thought for Detecting Fake News and Social Media DisinformationArkadiusz Modzelewski, Witold Sosnowski, Tiziano Labruna, Adam Wierzbicki et al.ACL 2025 · 8 citations
- Improving Task-Specific Multimodal Sentiment Analysis with General MLLMs via PromptingHaoyu Zhang, Yinan Zhang, Chaolong Ying, Xiaoying Tang et al.NeurIPS 2025 · 2 citations
- Exploring Explanations Improves the Robustness of In-Context LearningUkyo Honda, Tatsushi OkaACL 2025
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 272 citations
- Post Hoc Explanations of Language Models Can Improve Language ModelsSatyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun et al.NeurIPS 2023 · 87 citations
- One Prompt Word is Enough to Boost Adversarial Robustness for Pre-Trained Vision-Language ModelsLin Li, Haoyan Guan, Jianing Qiu, Michael W. SpratlingCVPR 2024
- Prompt Optimization via Adversarial In-Context LearningDo Xuan Long, Yiran Zhao, Hannah Brown, Yuxi Xie et al.ACL 2024 · 5 citations
- Same Question, Different Words: A Latent Adversarial Framework for Prompt RobustnessTingchen Fu, Fazl BarezEMNLP 2025
