OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints
Mingjie Pan, Jiyao Zhang, Tianshu Wu, Yinghao Zhao, Wenlong Gao, Hao Dong
Abstract
3 AgiBot https://omnimanip.github.io "Press button with hammer" "Press green button" "Insert flower into vase" "Insert pen into holder" "Close drawer" "Recycle battery" "Pour tea into the cup" "Pick up cup on dish" "Open drawer" "Put the lid on the teapot" "Close lid of laptop" "Open bottle" Passive Active Pour tea into the cup. Figure 1 . We proposed OmniManip, an open-vocabulary manipulation method that bridges the gap between the high-level reasoning of vision-language models (VLM) and the low-level precision, featuring closed-loop capabilities in both planning and execution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingYifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang et al.AAAI 2026 · 89 citations
- Exploring the Limits of Vision-Language-Action Manipulation in Cross-task GeneralizationJiaming Zhou, Ke Ye, Jiayi Liu, Teli Ma et al.NeurIPS 2025 · 43 citations
- Geometrically-Constrained Agent for Spatial ReasoningZeren Chen, Xiaoya Lu, Zhijie Zheng, Pengrui Li et al.CVPR 2026 · 29 citations
- HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM ReasoningZhi Jing, Siyuan Yang, Jicong Ao, Ting Xiao et al.NeurIPS 2025 · 23 citations
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous DrivingJianhua Han, Meng Tian, Jiangtong Zhu, Fan He et al.CVPR 2026 · 10 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic ManipulationKaidong Zhang, Rongtao Xu, Pengzhen Ren, Junfan Lin et al.ICCV 2025 · 1 citation
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin et al.AAAI 2026
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic ManipulationZhihao Zhu, Yifan Zheng, Siyu Pan, Yaohui Jin et al.ICCV 2025
- UniHM: Unified Dexterous Hand Manipulation with Vision Language ModelZhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya WangICLR 2026 · 4 citations
- Rethinking Intermediate Representation for VLM-based Robot ManipulationWeiliang Tang, Jialin Gao, Jia-Hui Pan, Gang Wang et al.CVPR 2026 · 2 citations
