CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based Agents
Huanxi Liu, Kun Hu, Qiang Wang, Yuanzhao Zhai, Feng Dawei, Bo Ding, Huaimin Wang
Abstract
Fine-tuning Large Language Models (LLMs) as autonomous agents on domain-specific data has emerged as a promising paradigm for tackling interactive, real-world tasks. However, existing studies have overlooked the critical coordination between long-term planning and multi-step execution in optimizing agent capabilities. This oversight leads to the propagation of impractical plans and plan-deviated trajectories within the optimization process, resulting in suboptimal task performance and hindering the further development of LLM-based agents in long-horizon tasks. To bridge this gap, we propose , a novel framework that explicitly integrates planning–execution coordination into LLM-based agent optimization. CoPE employs Self-Refining MCTS to generate task plans and multiple execution trajectories through environment interactions. By quantifying the coordination between planning and execution, CoPE assigns higher optimization weights to well-coordinated samples, enabling LLM-based agents to learn better planning and execution policies. Extensive experiments demonstrate that CoPE substantially improves agent coordination, outperforming state-of-the-art baselines on benchmarks comprising two long-horizon multi-step tasks. Codes and data are available at https://github.com/Octobrist/CoPE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc293176-d494-4466-9ffe-a16dcdc2bff2Builds on21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
Related papers
- CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent CooperationJie Liu, Pan Zhou, Yingjun Du, Ah-Hwee Tan et al.ICLR 2025
- Code Driven Planning with Domain-Adaptive SelectorZikang Tian, Shaohui Peng, Di Huang, Jiaming Guo et al.ICLR 2026
- Reflective Multi-Agent Collaboration based on Large Language ModelsXiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng et al.NeurIPS 2024 · 87 citations
- FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language AgentsQizheng Li, Yifei Zhang, Xiao Yang, Xu Yang et al.ICML 2026 · 3 citations
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu et al.NeurIPS 2024 · 54 citations
