Counterfactual Planning for Generalizable Agents' Actions
Jiarun Fu, Lizhong Ding, Qiuning Wei, Yuhan Guo, Yurong Cheng, Junyu Zhang
摘要
Large language models have revolutionized agent planning by serving as the engine of heuristic guidance. However, LLM-based agents often struggle to generalize across complex environments and to adapt to stochastic feedback arising from environment–action interactions. We propose Counterfactual Planning—a method designed to improve the generalizability and adaptability of agents' actions by inferring causal representations of environmental confounders and performing counterfactual reasoning over planned actions. We formalize the agent planning process as a structural causal model, providing a mathematical formulation for causal analysis of how environmental states influence action generation and how actions affect future state transitions. To support generalizable action planning, we introduce the State Causality Evaluator (SCE), which dynamically infers task-conditioned causal representations from complex environment states; and to enhance adaptability under stochastic feedback, we propose the What-If-Not (WIN) reward, which performs counterfactual interventions to refine actions through causal evaluation. We validate our framework in an open-world environment, where experiments demonstrate improvements in both action generalization and planning adaptability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FlowMAP: Flow Matching for Generalizable Agent PlanningJiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei 等ICML 2026
- Learning for Highly Faithful ExplainabilityYuhan Guo, Lizhong Ding, Shihao Jia, Yanyu Ren 等ICLR 2026
它引用的顶会 Paper8
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 被引用 193 次
- Foundation Models for Causal Inference via Prior-Data Fitted NetworksYuchen Ma, Dennis Frauen, Emil Javurek, Stefan FeuerriegelICLR 2026 · 被引用 37 次
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma 等ICLR 2024 · 被引用 31 次
- Curious Replay for Model-based AdaptationIsaac Kauvar, Chris Doyle, Linqi Zhou, Nick HaberICML 2023 · 被引用 18 次
相关 Paper
- Language Agents Meet Causality - Bridging LLMs and Causal World ModelsJohn Gkountouras, Matthias Lindemann, Phillip Lippe, Efstratios Gavves 等ICLR 2025
- Should I Have Expressed a Different Intent? Counterfactual Generation for LLM-Based Autonomous ControlAmirmohammad Farzaneh, Salvatore D'oro, Osvaldo SimeoneICML 2026
- Abstract Counterfactuals for Language Model AgentsEdoardo Pona, Milad Kazemi, Yali Du, David Watson 等NeurIPS 2025 · 被引用 3 次
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
- On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional StudyShuai Yang, Qi Yang, Luoxi Tang, Yuqiao Meng 等ICLR 2026 · 被引用 9 次
