ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
Kaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang, Zeju Li, Bolin Zhang, Wancai Zheng, Xinyi Yu, Hao Chen, Chunhua Shen
摘要
Language-guided long-horizon mobile manipulation has long been a grand challenge in embodied semantic reasoning, generalizable manipulation, and adaptive locomotion. Three fundamental limitations hinder progress: First, although large language models have shown promise in enhancing spatial reasoning and task planning through learned semantic priors, existing implementations remain confined to tabletop scenarios, failing to address the constrained perception and limited actuation ranges characteristic of mobile platforms. Second, current manipulation strategies exhibit insufficient generalization when confronted with the diverse object configurations encountered in open-world environments. Third, while crucial for practical deployment, the dual requirement of maintaining high platform maneuverability alongside precise end-effector control in unstructured settings remains understudied in the literature.
In this work, we present ODYSSEY, a unified mobile manipulation framework for agile quadruped robots equipped with manipulators, which seamlessly integrates high-level task planning with low-level whole-body control. To address the challenge of egocentric perception in language-conditioned tasks, we introduce a hierarchical planner powered by a vision-language model, enabling long-horizon instruction decomposition and precise action execution. At the control level, our novel whole-body policy achieves robust coordination of locomotion and manipulation across challenging terrains. We further present the first comprehensive benchmark for long-horizon mobile manipulation, evaluating diverse indoor and outdoor scenarios. Through successful sim-to-real transfer, we demonstrate the system’s generalization and robustness in real-world deployments, underscoring the practicality of legged manipulators in unstructured environments. Our work advances the feasibility of generalized robotic assistants capable of complex, dynamic tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State RepresentationMingyu Liu, Jiuhe Shu, Hui Chen, Zeju Li 等CVPR 2026 · 被引用 14 次
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic CameraHao Shi, Ze Wang, Shangwei Guo, Mengfei Duan 等CVPR 2026 · 被引用 11 次
- Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene GraphsJianing Qian, Qinhe Peng, Emmanuel Panov, Leonor Fermoselle 等CVPR 2026
- GAE: Unleashing Physical Potential of VLM with Generalizable Action ExpertMingyu Liu, Zheng Huang, Xiaoyi Lin, Muzhi Zhu 等ICML 2026
它引用的顶会 Paper5
- ARNOLD: A Benchmark for Language-Grounded Task Learning With Continuous States in Realistic 3D ScenesRan Gong, Jiangyong Huang, Yizhou Zhao, Haoran Geng 等ICCV 2023 · 被引用 77 次
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong 等NeurIPS 2025 · 被引用 65 次
- Open-Set Image Tagging with Multi-Grained Text SupervisionXinyu Huang, Yi-Jie Huang, Youcai Zhang, Weiwei Tian 等ACM MM 2025 · 被引用 13 次
- VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action TokenizersYating Wang, Haoyi Zhu, Mingyu Liu, Jiange Yang 等ICCV 2025 · 被引用 5 次
- OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial ConstraintsMingjie Pan, Jiyao Zhang, Tianshu Wu, Yinghao Zhao 等CVPR 2025
相关 Paper
- Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationFangyuan Wang, Shipeng Lyu, Peng Zhou, Anqing Duan 等AAAI 2025 · 被引用 9 次
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin 等AAAI 2026
- UniHM: Unified Dexterous Hand Manipulation with Vision Language ModelZhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya WangICLR 2026 · 被引用 4 次
- AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent PlanningZixuan Chen, Xiangrong Feng, Jieqi Shi, Lin Shao 等CVPR 2026
- AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile ManipulationSixiang Chen, Jiaming Liu, Siyuan Qian, Han Jiang 等NeurIPS 2025 · 被引用 30 次
