Discriminator-Guided Embodied Planning for LLM Agent
Haofu Qian, Chenjia Bai, Jiatao Zhang, Fei Wu, Wei Song, Xuelong Li
Abstract
Large Language Models (LLMs) have showcased remarkable reasoning capabilities in various domains, yet face challenges in complex embodied tasks due to coherent long-term policy, context-sensitive environmental understanding. Previous work performed LLM refinement relying on outcome-supervised feedback, which can be costly and ineffective. In this work, we introduce a novel framework, Discriminator-Guided Action OPtimization (DGAP) for facilitating optimization of LLM action plans via step-wise signals. Specifically, we employ a limited set of demonstrations to enable the discriminator in learning a score function, which assesses the alignment between LLM-generated action and the underlying optimal one at every step. Based on the discriminator, LLM is prompted to generate actions to maximize the score utilizing historical action-score pairs trajectory as guidance. Under mild conditions, DGAP resembles the critic-regularized optimization and is demonstrated to achieve a stronger policy than the LLM planner. In experiments across different LLMs (GPT-4, Llama3-70B) in ScienceWorld and VirtualHome, our method obtains superior performance and better efficiency than previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf630d07-6459-4329-a0b7-9bc276313b91Cited by top-tier papers2
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided ImprovementRuihan Yang, Fanghua Ye, Jian Li, Siyu Yuan et al.NeurIPS 2025 · 21 citations
- Counterfactual Planning for Generalizable Agents' ActionsJiarun Fu, Lizhong Ding, Qiuning Wei, Yuhan Guo et al.AAAI 2026
Builds on37
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Enhancing Decision-Making of Large Language Models via Actor-CriticHeng Dong, Kefei Duan, Chongjie ZhangICML 2025
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement LearningQihao Liu, Luoxin Ye, Wufei Ma, Yu-Cheng Chou et al.ICLR 2026 · 5 citations
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong et al.EMNLP 2023 · 109 citations
- Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationFangyuan Wang, Shipeng Lyu, Peng Zhou, Anqing Duan et al.AAAI 2025 · 9 citations
- GPO: Learning from Critical Steps to Improve LLM ReasoningJiahao Yu, Zelei Cheng, Xian Wu, Xinyu XingNeurIPS 2025 · 10 citations
