LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
Harry Mayne, Ryan Othniel Kearns, Yushi Yang, Andrew M. Bean, Eoin D. Delaney, Chris Russell, Adam Mahdi
Abstract
To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have predicted a different outcome. We evaluate whether LLMs can produce SCEs that are valid, achieving the intended outcome, and minimal, modifying the input no more than necessary. When asked to generate counterfactuals, we find that LLMs typically produce SCEs that are valid, but far from minimal, offering little insight into their decision-making behaviour. Worryingly, when asked to generate minimal counterfactuals, LLMs typically make excessively small edits that fail to change predictions. The observed validity-minimality trade-off is consistent across several LLMs, datasets, and evaluation settings. Our findings suggest that SCEs are, at best, an ineffective explainability tool and, at worst, can provide misleading insights into model behaviour. Proposals to deploy LLMs in high-stakes settings must consider the impact of unreliable self-explanations on downstream decision-making. Our code is available at https://github.com/HarryMayne/SCEs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f627c24-5cd1-48ef-bc0f-0243a4b832e8Cited by top-tier papers2
- Biases in the Blind Spot: Detecting What LLMs Fail to MentionIván Arcuschin, David Chanin, Adrià Garriga-Alonso, Oana-Maria CamburuICML 2026 · 7 citations
- What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input BeliefsNhi Nguyen, Shauli Ravfogel, Rajesh RanganathICML 2026
Builds on9
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- Measuring Association Between Labels and Free-Text RationalesSarah Wiegreffe, Ana Marasovic, Noah A. SmithEMNLP 2021 · 12 citations
Related papers
- Can LLMs Explain Themselves Counterfactually?Zahra Dehghanighobadi, Asja Fischer, Muhammad Bilal ZafarEMNLP 2025 · 1 citation
- Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric AssessmentsMarharyta Domnich, Julius Välja, Rasmus Moorits Veski, Giacomo Magnifico et al.AAAI 2025 · 11 citations
- Walk the Talk? Measuring the Faithfulness of Large Language Model ExplanationsKatie Matton, Robert Osazuwa Ness, John V. Guttag, Emre KicimanICLR 2025
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 11 citations
- CLOMO: Counterfactual Logical Modification with Large Language ModelsYinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao et al.ACL 2024 · 2 citations
