RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
Haifeng Huang, Xinyi Chen, Yilun Chen, Hao Li, Xiaoshen Han, Zehan Wang, Tai Wang, Jiangmiao Pang, Zhou Zhao
Abstract
Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing two key advantages: (1) effective spatial guidance that specifies target objects and placement areas while also conveying information about object shape and size, and (2) broad generalization potential driven by largescale vision-language models pretrained on diverse grounding datasets. We introduce ROBOGROUND, a groundingaware robotic manipulation system that leverages grounding masks as an intermediate representation to guide policy networks in object manipulation tasks. To further explore and enhance generalization, we propose an automated pipeline for generating large-scale, simulated data with a diverse set of objects and instructions. Extensive experiments show the value of our dataset and the effectiveness of grounding masks as intermediate guidance, significantly enhancing the generalization abilities of robot policies. Code and data are available at robo-ground.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and UnderstandingJunliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie et al.NeurIPS 2025 · 42 citations
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot PerceiverWenxuan Song, Ziyang Zhou, Han Zhao, Jiayi Chen et al.AAAI 2026 · 36 citations
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang et al.ICLR 2026 · 17 citations
- Spatially Guided Training for Vision-Language-Action ModelJinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu et al.ICLR 2026 · 6 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers et al.ICLR 2024 · 20 citations
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for RoboticsChan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree et al.CVPR 2025
- Bridging Scale Discrepancies in Robotic Control via Language-Based Action RepresentationsYuchi Zhang, Churui Sun, Shiqi Liang, Diyuan Liu et al.AAAI 2026
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory SketchesJiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu et al.ICLR 2024 · 135 citations
- ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and ReasoningZhenyang Liu, Yikai Wang, Sixiao Zheng, Tongying Pan et al.CVPR 2025
