World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
Siyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang, Panpan Cai, Jinlan Fu, Xipeng Qiu
Abstract
Recent advances in large vision-language models (LVLMs) have shown promise for embodied task planning, yet they struggle with fundamental challenges like dependency constraints and efficiency. Existing approaches either solely optimize action selection or leverage world models during inference, overlooking the benefits of learning to model the world as a way to enhance planning capabilities. We propose Dual Preference Optimization (DPO), a new learning framework that jointly optimizes state prediction and action selection through preference learning, enabling LVLMs to understand environment dynamics for better planning. To automatically collect trajectories and stepwise preference data without human annotation, we introduce a tree search mechanism for extensive exploration via trial-and-error. Extensive experiments on VoTa-Bench demonstrate that our DPO-based method significantly outperforms existing methods and GPT-4o when applied to Qwen2-VL (7B), LLaVA-1.6 (7B), and LLaMA-3.2 (11B), achieving superior task success rates with more efficient execution paths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60898680-9a18-458e-909d-e4b5d8a9d5fcCited by top-tier papers8
- From Word to World: Can Large Language Models be Implicit Text-based World Models?Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin et al.ACL 2026 · 27 citations
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied AgentsWanxin Tian, Shijie Zhang, Kevin Zhang, Xiaowei Chi et al.NeurIPS 2025 · 20 citations
- World-aware Planning Narratives Enhance Large Vision-Language Model PlannerJunhao Shi, Zhaoye Fei, Siyin Wang, Qipeng Guo et al.NeurIPS 2025 · 11 citations
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger LearningQiusi Zhan, Hyeonjeong Ha, Rui Yang, Sirui Xu et al.ICLR 2026 · 7 citations
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 6 citations
Builds on23
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
Related papers
- Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task AlignmentZiang Yan, Zhilin Li, Yinan He, Chenting Wang et al.CVPR 2025
- Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference TreesSijia Chen, Yibo Wang, Yi-Feng Wu, Qingguo Chen et al.NeurIPS 2024 · 51 citations
- VLP: Vision-Language Preference Learning for Embodied ManipulationRunze Liu, Chenjia Bai, Jiafei Lyu, Shengjie Sun et al.EMNLP 2025 · 1 citation
- Learning Planning-based Reasoning by Trajectories Collection and Process Reward SynthesizingFangkai Jiao, Chengwei Qin, Zhengyuan Liu, Nancy F. Chen et al.EMNLP 2024 · 1 citation
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang et al.ICML 2024 · 303 citations
