Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment
Shaofei Cai, Zhancun Mu, Anji Liu, Yitao Liang
摘要
We aim to develop a goal specification method that is semantically clear, spatially sensitive, domainagnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal alignment framework that allows users to specify target objects using segmentation masks from their camera views rather than the agent's observations. We highlight that behavior cloning alone fails to align the agent's behavior with human intent when the human and agent camera views differ significantly. To address this, we introduce two auxiliary objectives: cross-view consistency loss and target visibility loss, which explicitly enhance the agent's spatial reasoning ability. According to this, we develop RO C K E T-2, a state-of-the-art agent trained in Minecraft, achieving an improvement in the efficiency of inference 3× to 6× compared to ROCKET-1. We show that ROCKET-2 can directly interpret goals from human camera views, enabling better human-agent interaction. Remarkably, ROCKET-2 demonstrates zero-shot generalization capabilities: despite being trained exclusively on the Minecraft dataset, it can adapt and generalize to other 3D environments like Doom, DMLab, and Unreal through a simple action space mapping. The project page is available at https://craftjarvis.github. io/ROCKET-2/ . trade build a bridge use a portal activate ender portal shoot dragon rescue combat Minecraft Minecraft Minecraft Minecraft Minecraft Unreal Doom Figure 1 | Powered by cross-view goal specification, we are the first to show that AI agents can complete complex tasks such as building a bridge and damaging the dragon in Minecraft. In addition, it demonstrates an impressive zero-shot generalization to other 3D games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
相关 Paper
- ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context PromptingShaofei Cai, Zihao Wang, Kewei Lian, Zhancun Mu 等CVPR 2025
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma 等ICLR 2024 · 被引用 43 次
- Open-World Multi-Task Control Through Goal-Aware Representation Learning and Adaptive Horizon PredictionShaofei Cai, Zihao Wang, Xiaojian Ma, Anji Liu 等CVPR 2023
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement LearningKaichen He, Zihao Wang, Muyao Li, Anji Liu 等CVPR 2026
- Program Guided AgentShao-Hua Sun, Te-Lin Wu, Joseph J. LimICLR 2020 · 被引用 63 次
