Gumbel Counterfactual Generation From Language Models
Shauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson, Ryan Cotterell
摘要
Understanding and manipulating the causal generation mechanisms in language models is essential for controlling their behavior. Previous work has primarily relied on techniques such as representation surgery-e.g., model ablations or manipulation of linear subspaces tied to specific concepts-to intervene on these models. To understand the impact of interventions precisely, it is useful to examine counterfactuals-e.g., how a given sentence would have appeared had it been generated by the model following a specific intervention. We highlight that counterfactual reasoning is conceptually distinct from interventions, as articulated in Pearl's causal hierarchy. Based on this observation, we propose a framework for generating true string counterfactuals by reformulating language models as a structural equation model using the Gumbel-max trick, which we called Gumbel counterfactual generation. This reformulation allows us to model the joint distribution over original strings and their counterfactuals resulting from the same instantiation of the sampling noise. We develop an algorithm based on hindsight Gumbel sampling that allows us to infer the latent noise variables and generate counterfactuals of observed strings. Our experiments demonstrate that the approach produces meaningful counterfactuals while at the same time showing that commonly used intervention techniques have considerable undesired side effects. https://github.com/shauli-ravfogel/lm-counterfactuals * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Robust Reward Modeling via Causal RubricsPragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil 等ICLR 2026 · 被引用 20 次
- On the Reasoning Abilities of Masked Diffusion Language ModelsAnej Svete, Ashish SabharwalICLR 2026 · 被引用 8 次
- Counterfactual reasoning: an analysis of in-context emergenceMoritz Miller, Bernhard Schölkopf, Siyuan GuoNeurIPS 2025 · 被引用 5 次
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 被引用 4 次
- Abstract Counterfactuals for Language Model AgentsEdoardo Pona, Milad Kazemi, Yali Du, David Watson 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper19
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for TextNishtha Madaan, Inkit Padhi, Naveen Panwar, Diptikalyan SahaAAAI 2021 · 被引用 115 次
相关 Paper
- Learning Generalized Gumbel-max Causal MechanismsGuy Lorberbom, Daniel D. Johnson, Chris J. Maddison, Daniel Tarlow 等NeurIPS 2021 · 被引用 25 次
- Counterfactual Temporal Point ProcessesKimia Noorbakhsh, Manuel Gomez-RodriguezNeurIPS 2022 · 被引用 31 次
- From Probability to Counterfactuals: the Increasing Complexity of Satisfiability in Pearl's Causal HierarchyJulian Dörfler, Benito van der Zander, Markus Bläser, Maciej LiskiewiczICLR 2025 · 被引用 1 次
- Natural Counterfactuals With Necessary BacktrackingGuang-Yuan Hao, Jiji Zhang, Biwei Huang, Hao Wang 等NeurIPS 2024 · 被引用 2 次
- Counterfactual Explanations in Sequential Decision Making Under UncertaintyStratis Tsirtsis, Abir De, Manuel Gomez RodriguezNeurIPS 2021 · 被引用 59 次
