Abstract Counterfactuals for Language Model Agents
Edoardo Pona, Milad Kazemi, Yali Du, David Watson, Nicola Paoletti
摘要
Counterfactual inference is a powerful tool for analysing and evaluating autonomous agents, but its application to language model (LM) agents remains challenging. Existing work on counterfactuals in LMs has primarily focused on token-level counterfactuals, which are often inadequate for LM agents due to their open-ended action spaces. Unlike traditional agents with fixed, clearly defined action spaces, the actions of LM agents are often implicit in the strings they output, making their action spaces difficult to define and interpret. Furthermore, the meanings of individual tokens can shift depending on the context, adding complexity to token-level reasoning and sometimes leading to biased or meaningless counterfactuals. We introduce Abstract Counterfactuals, a framework that emphasises high-level characteristics of actions and interactions within an environment, enabling counterfactual reasoning tailored to user-relevant features. Our experiments demonstrate that the approach produces consistent and meaningful counterfactuals while minimising the undesired side effects of token-level methods. We conduct experiments on text-based games and counterfactual text generation, while considering both token-level and latent-space interventions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Representation Surgery: Theory and Practice of Affine SteeringShashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni 等ICML 2024 · 被引用 36 次
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen 等ACL 2020 · 被引用 16 次
- Gumbel Counterfactual Generation From Language ModelsShauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson, Ryan CotterellICLR 2025
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
相关 Paper
- Counterfactual Planning for Generalizable Agents' ActionsJiarun Fu, Lizhong Ding, Qiuning Wei, Yuhan Guo 等AAAI 2026
- Should I Have Expressed a Different Intent? Counterfactual Generation for LLM-Based Autonomous ControlAmirmohammad Farzaneh, Salvatore D'oro, Osvaldo SimeoneICML 2026
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell 等EMNLP 2023 · 被引用 57 次
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira 等EMNLP 2023 · 被引用 6 次
- Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy OptimizationZelai Xu, Wanjun Gu, Chao Yu, Yi Wu 等ICML 2025
