Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment
Shaofei Cai, Zhancun Mu, Anji Liu, Yitao Liang
Abstract
We aim to develop a goal specification method that is semantically clear, spatially sensitive, domainagnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal alignment framework that allows users to specify target objects using segmentation masks from their camera views rather than the agent's observations. We highlight that behavior cloning alone fails to align the agent's behavior with human intent when the human and agent camera views differ significantly. To address this, we introduce two auxiliary objectives: cross-view consistency loss and target visibility loss, which explicitly enhance the agent's spatial reasoning ability. According to this, we develop RO C K E T-2, a state-of-the-art agent trained in Minecraft, achieving an improvement in the efficiency of inference 3× to 6× compared to ROCKET-1. We show that ROCKET-2 can directly interpret goals from human camera views, enabling better human-agent interaction. Remarkably, ROCKET-2 demonstrates zero-shot generalization capabilities: despite being trained exclusively on the Minecraft dataset, it can adapt and generalize to other 3D environments like Doom, DMLab, and Unreal through a simple action space mapping. The project page is available at https://craftjarvis.github. io/ROCKET-2/ . trade build a bridge use a portal activate ender portal shoot dragon rescue combat Minecraft Minecraft Minecraft Minecraft Minecraft Unreal Doom Figure 1 | Powered by cross-view goal specification, we are the first to show that AI agents can complete complex tasks such as building a bridge and damaging the dragon in Minecraft. In addition, it demonstrates an impressive zero-shot generalization to other 3D games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ab8aef5-dd94-44f1-8967-24dc952a74baBuilds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
Related papers
- ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context PromptingShaofei Cai, Zihao Wang, Kewei Lian, Zhancun Mu et al.CVPR 2025
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma et al.ICLR 2024 · 43 citations
- Open-World Multi-Task Control Through Goal-Aware Representation Learning and Adaptive Horizon PredictionShaofei Cai, Zihao Wang, Xiaojian Ma, Anji Liu et al.CVPR 2023
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement LearningKaichen He, Zihao Wang, Muyao Li, Anji Liu et al.CVPR 2026
- Program Guided AgentShao-Hua Sun, Te-Lin Wu, Joseph J. LimICLR 2020 · 63 citations
