Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators
Qiwei Liang, Boyang Cai, Rongyi He, Hui Li, Tao Teng, Haihan Duan, Changxin Huang, Runhao Zeng
Abstract
Quadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, existing research has predominantly focused on static-object grasping, neglecting the challenges posed by dynamic targets and thus limiting applicability in dynamic scenarios such as logistics sorting and human–robot collaboration. To address this, we introduce DQ-Bench, a new benchmark that systematically evaluates dynamic grasping across varying object motions, velocities, heights, object types, and terrain complexities, along with comprehensive evaluation metrics. Building upon this benchmark, we propose DQ-Net, a compact teacher–student framework designed to infer grasp configurations from limited perceptual cues. During training, the teacher network leverages privileged information to holistically model both the static geometric properties and dynamic motion characteristics of the target, and integrates a grasp fusion module to deliver robust guidance for motion planning. Concurrently, we design a lightweight student network that performs dual-viewpoint temporal modeling using only the target mask, depth map, and proprioceptive state, enabling closed-loop action outputs without reliance on privileged data. Extensive experiments on DQ-Bench demonstrate that DQ-Net achieves robust dynamic objects grasping across multiple task settings, substantially outperforming baseline methods in both success rate and responsiveness. We will release our codebase and benchmark publicly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ec93277-4589-42fc-96a5-ae50f9a86361Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersRuihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu et al.ICLR 2022 · 146 citations
- Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot ResponseJunfeng Long, Zirui Wang, Quanyi Li, Liu Cao et al.ICLR 2024 · 66 citations
- GraspNet-1Billion: A Large-Scale Benchmark for General Object GraspingHaoshu Fang, Chenxi Wang, Minghao Gou, Cewu LuCVPR 2020
- Target-referenced Reactive Grasping for Dynamic ObjectsJirong Liu, Ruo Zhang, Haoshu Fang, Minghao Gou et al.CVPR 2023
Related papers
- DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-To-Robot HandoverYouzhuo Wang, Jiayi Ye, Chuyang Xiao, Yiming Zhong et al.ICCV 2025 · 1 citation
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D WorldYuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu et al.CVPR 2026 · 15 citations
- RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General GraspingDongming Wu, Yanping Fu, Saike Huang, Yingfei Liu et al.ICCV 2025 · 2 citations
- Towards Diverse Behaviors: A Benchmark for Imitation Learning with Human DemonstrationsXiaogang Jia, Denis Blessing, Xinkai Jiang, Moritz Reuss et al.ICLR 2024 · 48 citations
- Generalizing 6-DoF Grasp Detection via Domain Prior KnowledgeHaoxiang Ma, Modi Shi, Boyang Gao, Di HuangCVPR 2024 · 10 citations
