Abstract Counterfactuals for Language Model Agents
Edoardo Pona, Milad Kazemi, Yali Du, David Watson, Nicola Paoletti
Abstract
Counterfactual inference is a powerful tool for analysing and evaluating autonomous agents, but its application to language model (LM) agents remains challenging. Existing work on counterfactuals in LMs has primarily focused on token-level counterfactuals, which are often inadequate for LM agents due to their open-ended action spaces. Unlike traditional agents with fixed, clearly defined action spaces, the actions of LM agents are often implicit in the strings they output, making their action spaces difficult to define and interpret. Furthermore, the meanings of individual tokens can shift depending on the context, adding complexity to token-level reasoning and sometimes leading to biased or meaningless counterfactuals. We introduce Abstract Counterfactuals, a framework that emphasises high-level characteristics of actions and interactions within an environment, enabling counterfactual reasoning tailored to user-relevant features. Our experiments demonstrate that the approach produces consistent and meaningful counterfactuals while minimising the undesired side effects of token-level methods. We conduct experiments on text-based games and counterfactual text generation, while considering both token-level and latent-space interventions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85983585-b951-4f5f-99c8-fac7a895ee19Builds on5
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Representation Surgery: Theory and Practice of Affine SteeringShashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni et al.ICML 2024 · 36 citations
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen et al.ACL 2020 · 16 citations
- Gumbel Counterfactual Generation From Language ModelsShauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson, Ryan CotterellICLR 2025
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
Related papers
- Counterfactual Planning for Generalizable Agents' ActionsJiarun Fu, Lizhong Ding, Qiuning Wei, Yuhan Guo et al.AAAI 2026
- Should I Have Expressed a Different Intent? Counterfactual Generation for LLM-Based Autonomous ControlAmirmohammad Farzaneh, Salvatore D'oro, Osvaldo SimeoneICML 2026
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell et al.EMNLP 2023 · 57 citations
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira et al.EMNLP 2023 · 6 citations
- Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy OptimizationZelai Xu, Wanjun Gu, Chao Yu, Yi Wu et al.ICML 2025
