GROOT-2: Weakly Supervised Multimodal Instruction Following Agents
Shaofei Cai, Bowei Zhang, Zihao Wang, Haowei Lin, Xiaojian Ma, Anji Liu, Yitao Liang
Abstract
Developing agents that can follow multimodal instructions remains a fundamental challenge in robotics and AI. Although large-scale pre-training on unlabeled datasets (no language instruction) has enabled agents to learn diverse behaviors, these agents often struggle with following instructions. While augmenting the dataset with instruction labels can mitigate this issue, acquiring such high-quality annotations at scale is impractical. To address this issue, we frame the problem as a semi-supervised learning task and introduce GROOT-2, a multimodal instructable agent trained using a novel approach that combines weak supervision with latent variable models. Our method consists of two key components: constrained selfimitating, which utilizes large amounts of unlabeled demonstrations to enable the policy to learn diverse behaviors, and human intention alignment, which uses a smaller set of labeled demonstrations to ensure the latent space reflects human intentions. GROOT-2's effectiveness is validated across four diverse environments, ranging from video games to robotic manipulation, demonstrating its robust multimodal instruction-following capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a3e50eb-ff79-4aca-9da6-789e5c984173Cited by top-tier papers8
- Preference Goal Tuning: Post-Training as Latent Control for Frozen PoliciesGuangyu Zhao, Kewei Lian, Haoxuan Ru, Borong Zhang et al.ICML 2026 · 5 citations
- Open-World Skill Discovery from Unsegmented Demonstration VideosJingwen Deng, Zihao Wang, Shaofei Cai, Anji Liu et al.ICCV 2025 · 5 citations
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian et al.CVPR 2026 · 4 citations
- World2Minecraft: Occupancy-Driven Simulated Scenes ConstructionLechao Zhang, Haoran Xu, Jingyu Gong, Xuhong Wang et al.ICLR 2026 · 1 citation
- Steering Visuomotor Policy in Open Worlds via Cross-View Goal AlignmentShaofei Cai, Zhancun Mu, Anji Liu, Yitao LiangAAAI 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
Related papers
- Ess-InfoGAIL: Semi-supervised Imitation Learning from Imbalanced DemonstrationsHuiqiao Fu, Kaiqiang Tang, Yuanyang Lu, Yiming Qi et al.NeurIPS 2023 · 15 citations
- Data Augmentation for Instruction Following Policies via Trajectory SegmentationNiklas Höpner, Ilaria Tiddi, Herke van HoofAAAI 2025
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent PretrainingWeimin Xiong, Shuhao Gu, Bowen Ye, Zihao Yue et al.ICML 2026 · 2 citations
- The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity DataThomas Pouplin, Kasia Kobalczyk, Hao Sun, Mihaela van der SchaarICML 2025
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang et al.ICML 2026 · 3 citations
