GROOT: Learning to Follow Instructions by Watching Gameplay Videos
Shaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma, Anji Liu, Yitao Liang
摘要
We study the problem of building a controller that can follow open-ended instructions in open-world environments. We propose to follow reference videos as instructions, which offer expressive goal specifications while eliminating the need for expensive text-gameplay annotations. A new learning framework is derived to allow learning such instruction-following controllers from gameplay videos while producing a video instruction encoder that induces a structured goal space. We implement our agent G ROO T in a simple yet effective encoder-decoder architecture based on causal transformers. We evaluate G ROO T against open-world counterparts and human players on a proposed Minecraft SkillForge benchmark. The Elo ratings clearly show that G ROO T is closing the human-machine gap as well as exhibiting a 70% winning rate over the best generalist agent baseline. Qualitative analysis of the induced goal space further demonstrates some interesting emergent properties, including the goal composition and complex gameplay behavior synthesis. The project page is available at https: //craftjarvis-groot.github.io. Figure 1 | Through the cultivation of extensive gameplay videos, G ROO T has grown a rich set of skill fruits (number denotes success rate; skills shown above do not mean to be exhaustive; kudos to our artist Haowei).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 等NeurIPS 2024 · 被引用 104 次
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in MinecraftZihao Wang, Muyao Li, Kaichen He, Xiangyu Wang 等ICML 2026 · 被引用 8 次
- Open-World Skill Discovery from Unsegmented Demonstration VideosJingwen Deng, Zihao Wang, Shaofei Cai, Anji Liu 等ICCV 2025 · 被引用 5 次
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian 等CVPR 2026 · 被引用 4 次
- GameVerse: Can Vision-Language Models Learn from Video-based Reflection?Kuan Zhang, Dongchen Liu, Qiyue Zhao, Jinkun Hou 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
相关 Paper
- GameFactorly: Creating New Games with Generative Interactive VideosJiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan 等ICCV 2025 · 被引用 7 次
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba 等NeurIPS 2023 · 被引用 123 次
- Steering Visuomotor Policy in Open Worlds via Cross-View Goal AlignmentShaofei Cai, Zhancun Mu, Anji Liu, Yitao LiangAAAI 2026
- OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following AgentsZihao Wang, Shaofei Cai, Zhancun Mu, Haowei Lin 等NeurIPS 2024 · 被引用 37 次
- Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned PolicyZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 等CVPR 2025
