TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making
Kechen Jiao, Zhirui Fang, Jiahao Liu, Bei Li, Qifan Wang, Xinyu Liu, Junhao Ruan, Zhongjian Qiao, Yifan Zhu, Yaxin Xu, Jingang Wang, Xiu Li
摘要
Using effective generalization capabilities of vision language models (VLMs) in contextspecific dynamic tasks for embodied artificial intelligence remains a significant challenge. Although supervised fine-tuned models can better align with the real physical world, they still exhibit sluggish responses and hallucination issues in dynamically changing environments, necessitating further alignment. Existing post-SFT methods, reliant on reinforcement learning and chain-of-thought (CoT) approaches, are constrained by sparse rewards and actiononly optimization, resulting in low sample efficiency, poor consistency, and model degradation. To address these issues, this paper proposes Thought-Centric Preference Optimization (TCPO) for effective embodied decisionmaking. Specifically, TCPO introduces a stepwise preference-based optimization approach, transforming sparse reward signals into richer step sample pairs. It emphasizes the alignment of the model's intermediate reasoning process, mitigating the problem of model degradation. Moreover, by incorporating Action Policy Consistency Constraint (APC), it further imposes consistency constraints on the model output. Experiments in the ALFWorld environment demonstrate an average success rate of 26.67%, achieving a 6% improvement over RL4VLM and validating the effectiveness of our approach in mitigating model degradation after fine-tuning. These results highlight the potential of integrating preferencebased learning techniques with CoT processes to enhance the decision-making capabilities of vision-language models in embodied agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference LearningHuang Kun, Weikai Xu, Yuxuan Liu, Quandong Wang 等ICLR 2026 · 被引用 9 次
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics ShiftsZhongjian Qiao, Rui Yang, Jiafei Lyu, Xiu Li 等ICLR 2026 · 被引用 7 次
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 被引用 6 次
- Unifying Value Alignment and Assignment in Cross-Domain Offline Reinforcement Learning with Heterogeneous DatasetsZhongjian Qiao, Jiafei Lyu, Chenjia Bai, Peisong Wang 等ICML 2026
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
相关 Paper
- GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksHaoqiang Kang, Enna Sachdeva, Piyush Gupta, Sangjae Bae 等CVPR 2025
- GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-Based VLM Agent TrainingTong Wei, Yijun Yang, Junliang Xing, Yuanchun Shi 等ICCV 2025 · 被引用 1 次
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied AgentsWanxin Tian, Shijie Zhang, Kevin Zhang, Xiaowei Chi 等NeurIPS 2025 · 被引用 20 次
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan 等NeurIPS 2024 · 被引用 214 次
- Model-Based Imaginative Planning for Embodied AgentsJunru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang 等ACL 2026
