RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
Mingxuan Yan, Yuping Wang, Zechun Liu, Jiachen Li
摘要
To tackle long-horizon tasks, recent hierarchical vision-language-action (VLAs) frameworks employ vision-language model (VLM)-based planners to decompose complex manipulation tasks into simpler sub-tasks that low-level visuomotor policies can handle. Typically, the VLM planner needs finetuning to learn to decompose a new task, which requires target task demonstrations segmented into sub-tasks by either human annotation or heuristic rules. However, without prior knowledge, the heuristic sub-tasks can deviate significantly from the visuomotor policy's training data, thereby degrading task performance. To address these issues, we propose a Retrieval-based Demonstration Decomposer (RDD) that automatically decomposes video demonstrations into sub-tasks with prior by aligning the visual features of the decomposed sub-task intervals with those from the training data of the low-level visuomotor policies. RDD outperforms the state-of-the-art sub-task decomposer on both simulation and real-world tasks, demonstrating robustness across diverse settings. Code and more results are available at rdd-neurips.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized DrivingZehao Wang, Huaide Jiang, Shuaiwu Dong, Yuping Wang 等CVPR 2026 · 被引用 7 次
- Motion Dynamics Learning for Few-Shot Embodied AdaptationSibo He, Weiying Xie, Daixun Li, Junhao Zhong 等ICML 2026
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 等NeurIPS 2023 · 被引用 336 次
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani 等ICML 2023 · 被引用 212 次
相关 Paper
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen 等AAAI 2026 · 被引用 1 次
- VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsGuangyan Chen, Meiling Wang, Te Cui, Yao Mu 等NeurIPS 2024 · 被引用 24 次
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 被引用 152 次
- KISA: A Unified Keyframe Identifier and Skill Annotator for Long-Horizon Robotics DemonstrationsLongxin Kou, Fei Ni, Yan Zheng, Jinyi Liu 等ICML 2024 · 被引用 5 次
- FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic ManipulationGanlong Zhao, Zijia Tang, Xingping Chen, Zhanghui Kuang 等CVPR 2026 · 被引用 10 次
