Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
Katie Matton, Robert Osazuwa Ness, John V. Guttag, Emre Kiciman
Abstract
Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be unfaithful. This, in turn, can lead to over-trust and misuse. We introduce a new approach for measuring the faithfulness of LLM explanations. First, we provide a rigorous definition of faithfulness. Since LLM explanations mimic human explanations, they often reference high-level concepts in the input question that purportedly influenced the model. We define faithfulness in terms of the difference between the set of concepts that the LLM's explanations imply are influential and the set that truly are. Second, we present a novel method for estimating faithfulness that is based on: (1) using an auxiliary LLM to modify the values of concepts within model inputs to create realistic counterfactuals, and (2) using a hierarchical Bayesian model to quantify the causal effects of concepts at both the example- and dataset-level. Our experiments show that our method can be used to quantify and discover interpretable patterns of unfaithfulness. On a social bias task, we uncover cases where LLM explanations hide the influence of social bias. On a medical question answering task, we uncover cases where LLM explanations provide misleading claims about which pieces of evidence influenced the model's decisions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 567bc231-cc72-45ec-b3a3-9111da95b559Cited by top-tier papers12
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought ReasoningXu Shen, Song Wang, Zhen Tan, Laura Yao et al.ICLR 2026 · 28 citations
- A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice ExperimentsManuel Cherep, Chengtian Ma, Abigail Xu, Maya Shaked et al.ICLR 2026 · 13 citations
- RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning ModelsYunseok Han, Yejoon Lee, Jaeyoung DoICLR 2026 · 10 citations
- How Does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse AutoencodingXi Chen, Aske Plaat, Niki van SteinAAAI 2026 · 9 citations
- A Positive Case for Faithfulness: Explanations Help Predict Model BehaviorHarry Mayne, Justin S. Kang, Dewi Gould, Kannan Ramchandran et al.ICML 2026 · 9 citations
Builds on6
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorEldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat et al.NeurIPS 2022 · 69 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksMaksym Andriushchenko, Francesco Croce, Nicolas FlammarionICLR 2025 · 7 citations
Related papers
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language ModelsWei Jie Yeo, Ranjan Satapathy, Erik CambriaEMNLP 2025 · 2 citations
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- A Causal Lens for Evaluating Faithfulness MetricsKerem Zaman, Shashank SrivastavaEMNLP 2025
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo et al.VLDB 2026 · 2 citations
