Rethinking Intermediate Representation for VLM-based Robot Manipulation
Weiliang Tang, Jialin Gao, Jia-Hui Pan, Gang Wang, Li Erran Li, Yun-Hui Liu, Mingyu Ding, Pheng-Ann Heng, Chi-Wing Fu
Abstract
Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free grammar, we design the Semantic Assembly representation named SEAM, by decomposing the intermediate representation into vocabulary and grammar. Doing so leads us to a concise vocabulary of semantically-rich operations and a VLM-friendly grammar for handling diverse unseen tasks. In addition, we design a new open-vocabulary segmentation paradigm with a retrieval-augmented few-shot learning strategy to localize fine-grained object parts for manipulation, effectively with the shortest inference time over all state-of-the-art parallel works. Also, we formulate new metrics for action-generalizability and VLM-comprehensibility, demonstrating the compelling performance of SEAM over mainstream representations on both aspects. Extensive real-world experiments further manifest its SOTA performance under varying settings and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1983317-8e51-42d6-9a50-c0fef2e86a0dBuilds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
Related papers
- Programmatically Grounded, Compositionally Generalizable Robotic ManipulationRenhao Wang, Jiayuan Mao, Joy Hsu, Hang Zhao et al.ICLR 2023 · 4 citations
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin et al.AAAI 2026
- Language-Grounded Decoupled Action Representation for Robotic ManipulationWuDing Weng, Tongshu Wu, Liucheng Chen, Siyu xie et al.CVPR 2026 · 2 citations
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic ManipulationZhihao Zhu, Yifan Zheng, Siyu Pan, Yaohui Jin et al.ICCV 2025
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language MappingDanyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang et al.ACM MM 2025 · 2 citations
