Lune

EMNLP2025Top-tier venue

Can LLMs Explain Themselves Counterfactually?

Zahra Dehghanighobadi, Asja Fischer, Muhammad Bilal Zafar

2025Year
1Citations
2Top-tier citations

Abstract

Explanations are an important tool for gaining insights into model behavior, calibrating user trust, and ensuring compliance.The past few years have seen a flurry of methods for generating explanations, many of which involve computing model gradients or solving specially designed optimization problems.Owing to the remarkable reasoning abilities of LLMs, selfexplanation, i.e., prompting the model to explain its outputs, has recently emerged as a new paradigm.We study a specific type of self-explanation, self-generated counterfactual explanations (SCEs).We test LLMs' ability to generate SCEs across families, sizes, temperatures, and datasets.We find that LLMs sometimes struggle to generate SCEs.When they do, their prediction often does not agree with their own counterfactual reasoning.github.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers2

Ask how each one uses it

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines