Can LLMs Explain Themselves Counterfactually?
Zahra Dehghanighobadi, Asja Fischer, Muhammad Bilal Zafar
摘要
Explanations are an important tool for gaining insights into model behavior, calibrating user trust, and ensuring compliance.The past few years have seen a flurry of methods for generating explanations, many of which involve computing model gradients or solving specially designed optimization problems.Owing to the remarkable reasoning abilities of LLMs, selfexplanation, i.e., prompting the model to explain its outputs, has recently emerged as a new paradigm.We study a specific type of self-explanation, self-generated counterfactual explanations (SCEs).We test LLMs' ability to generate SCEs across families, sizes, temperatures, and datasets.We find that LLMs sometimes struggle to generate SCEs.When they do, their prediction often does not agree with their own counterfactual reasoning.github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual ExplanationsHarry Mayne, Ryan Othniel Kearns, Yushi Yang, Andrew M. Bean 等EMNLP 2025 · 被引用 9 次
- Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and RecoverabilityAdvaith Malladi, Shashank SrivastavaICML 2026
它引用的顶会 Paper18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
相关 Paper
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
- Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric AssessmentsMarharyta Domnich, Julius Välja, Rasmus Moorits Veski, Giacomo Magnifico 等AAAI 2025 · 被引用 11 次
- SaySelf: Teaching LLMs to Express Confidence with Self-Reflective RationalesTianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu 等EMNLP 2024 · 被引用 10 次
- What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input BeliefsNhi Nguyen, Shauli Ravfogel, Rajesh RanganathICML 2026
- Are Human-generated Demonstrations Necessary for In-context Learning?Rui Li, Guoyin Wang, Jiwei LiICLR 2024 · 被引用 17 次
