Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
Katie Matton, Robert Osazuwa Ness, John V. Guttag, Emre Kiciman
摘要
Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be unfaithful. This, in turn, can lead to over-trust and misuse. We introduce a new approach for measuring the faithfulness of LLM explanations. First, we provide a rigorous definition of faithfulness. Since LLM explanations mimic human explanations, they often reference high-level concepts in the input question that purportedly influenced the model. We define faithfulness in terms of the difference between the set of concepts that the LLM's explanations imply are influential and the set that truly are. Second, we present a novel method for estimating faithfulness that is based on: (1) using an auxiliary LLM to modify the values of concepts within model inputs to create realistic counterfactuals, and (2) using a hierarchical Bayesian model to quantify the causal effects of concepts at both the example- and dataset-level. Our experiments show that our method can be used to quantify and discover interpretable patterns of unfaithfulness. On a social bias task, we uncover cases where LLM explanations hide the influence of social bias. On a medical question answering task, we uncover cases where LLM explanations provide misleading claims about which pieces of evidence influenced the model's decisions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought ReasoningXu Shen, Song Wang, Zhen Tan, Laura Yao 等ICLR 2026 · 被引用 28 次
- A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice ExperimentsManuel Cherep, Chengtian Ma, Abigail Xu, Maya Shaked 等ICLR 2026 · 被引用 13 次
- RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning ModelsYunseok Han, Yejoon Lee, Jaeyoung DoICLR 2026 · 被引用 10 次
- How Does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse AutoencodingXi Chen, Aske Plaat, Niki van SteinAAAI 2026 · 被引用 9 次
- A Positive Case for Faithfulness: Explanations Help Predict Model BehaviorHarry Mayne, Justin S. Kang, Dewi Gould, Kannan Ramchandran 等ICML 2026 · 被引用 9 次
它引用的顶会 Paper6
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
- CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorEldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat 等NeurIPS 2022 · 被引用 69 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksMaksym Andriushchenko, Francesco Croce, Nicolas FlammarionICLR 2025 · 被引用 7 次
相关 Paper
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language ModelsWei Jie Yeo, Ranjan Satapathy, Erik CambriaEMNLP 2025 · 被引用 2 次
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- A Causal Lens for Evaluating Faithfulness MetricsKerem Zaman, Shashank SrivastavaEMNLP 2025
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo 等VLDB 2026 · 被引用 2 次
