ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
Shaofei Cai, Zihao Wang, Kewei Lian, Zhancun Mu, Xiaojian Ma, Anji Liu, Yitao Liang
摘要
Abstract Vision-language models (VLMs) have excelled in multimodal tasks, but adapting them to embodied decision-making in open-world environments presents challenges. One critical issue is bridging the gap between discrete entities in lowlevel observations and the abstract concepts required for effective planning. A common solution is building hierarchical agents, where VLMs serve as high-level reasoners that break down tasks into executable sub-tasks, typically specified using language. However, language suffers from the inability to communicate detailed spatial information. We propose visual-temporal context prompting, a novel communication protocol between VLMs and policy models. This protocol leverages object segmentation from past observations to guide policy-environment interactions. Using this approach, we train ROCKET-1, a low-level policy that predicts actions based on concatenated visual observations and segmentation masks, supported by real-time object tracking from SAM-2. Our method unlocks the potential of VLMs, enabling them to tackle complex tasks that demand spatial reasoning. Experiments in Minecraft show that our approach enables agents to achieve previously unattainable tasks, with a 76% absolute improvement in open-world interaction performance. Codes are available at https://craftjarvis.github.io/ROCKET-1 . This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive TasksWenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu 等ACL 2026 · 被引用 53 次
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in MinecraftZihao Wang, Muyao Li, Kaichen He, Xiangyu Wang 等ICML 2026 · 被引用 8 次
- Preference Goal Tuning: Post-Training as Latent Control for Frozen PoliciesGuangyu Zhao, Kewei Lian, Haoxuan Ru, Borong Zhang 等ICML 2026 · 被引用 5 次
- World2Minecraft: Occupancy-Driven Simulated Scenes ConstructionLechao Zhang, Haoran Xu, Jingyu Gong, Xuhong Wang 等ICLR 2026 · 被引用 1 次
- Steering Visuomotor Policy in Open Worlds via Cross-View Goal AlignmentShaofei Cai, Zhancun Mu, Anji Liu, Yitao LiangAAAI 2026
它引用的顶会 Paper14
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman 等NeurIPS 2022 · 被引用 344 次
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu 等NeurIPS 2023 · 被引用 178 次
相关 Paper
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee 等NeurIPS 2025 · 被引用 15 次
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 被引用 6 次
- Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationJiaqi Chen, Bingqian Lin, Xinmin Liu, Lin Ma 等AAAI 2025 · 被引用 61 次
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied TasksPhilip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo 等NeurIPS 2025 · 被引用 8 次
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu 等AAAI 2024 · 被引用 122 次
