Gumbel Counterfactual Generation From Language Models
Shauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson, Ryan Cotterell
Abstract
Understanding and manipulating the causal generation mechanisms in language models is essential for controlling their behavior. Previous work has primarily relied on techniques such as representation surgery-e.g., model ablations or manipulation of linear subspaces tied to specific concepts-to intervene on these models. To understand the impact of interventions precisely, it is useful to examine counterfactuals-e.g., how a given sentence would have appeared had it been generated by the model following a specific intervention. We highlight that counterfactual reasoning is conceptually distinct from interventions, as articulated in Pearl's causal hierarchy. Based on this observation, we propose a framework for generating true string counterfactuals by reformulating language models as a structural equation model using the Gumbel-max trick, which we called Gumbel counterfactual generation. This reformulation allows us to model the joint distribution over original strings and their counterfactuals resulting from the same instantiation of the sampling noise. We develop an algorithm based on hindsight Gumbel sampling that allows us to infer the latent noise variables and generate counterfactuals of observed strings. Our experiments demonstrate that the approach produces meaningful counterfactuals while at the same time showing that commonly used intervention techniques have considerable undesired side effects. https://github.com/shauli-ravfogel/lm-counterfactuals * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d08e013-e412-4433-8877-b15bf1d67399Cited by top-tier papers8
- Robust Reward Modeling via Causal RubricsPragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil et al.ICLR 2026 · 20 citations
- On the Reasoning Abilities of Masked Diffusion Language ModelsAnej Svete, Ashish SabharwalICLR 2026 · 8 citations
- Counterfactual reasoning: an analysis of in-context emergenceMoritz Miller, Bernhard Schölkopf, Siyuan GuoNeurIPS 2025 · 5 citations
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 4 citations
- Abstract Counterfactuals for Language Model AgentsEdoardo Pona, Milad Kazemi, Yali Du, David Watson et al.NeurIPS 2025 · 3 citations
Builds on19
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for TextNishtha Madaan, Inkit Padhi, Naveen Panwar, Diptikalyan SahaAAAI 2021 · 115 citations
Related papers
- Learning Generalized Gumbel-max Causal MechanismsGuy Lorberbom, Daniel D. Johnson, Chris J. Maddison, Daniel Tarlow et al.NeurIPS 2021 · 25 citations
- Counterfactual Temporal Point ProcessesKimia Noorbakhsh, Manuel Gomez-RodriguezNeurIPS 2022 · 31 citations
- From Probability to Counterfactuals: the Increasing Complexity of Satisfiability in Pearl's Causal HierarchyJulian Dörfler, Benito van der Zander, Markus Bläser, Maciej LiskiewiczICLR 2025 · 1 citation
- Natural Counterfactuals With Necessary BacktrackingGuang-Yuan Hao, Jiji Zhang, Biwei Huang, Hao Wang et al.NeurIPS 2024 · 2 citations
- Counterfactual Explanations in Sequential Decision Making Under UncertaintyStratis Tsirtsis, Abir De, Manuel Gomez RodriguezNeurIPS 2021 · 59 citations
