Catch & Carry: reusable neural controllers for vision-guided whole-body tasks
Josh Merel, Saran Tunyasuvunakool, Arun Ahuja, Yuval Tassa, Leonard Hasenclever, Vu Pham, Tom Erez, Greg Wayne, Nicolas Heess
Abstract
Fig. 1. A catch-carry-toss sequence (bottom) from first-person visual inputs (top). Note how the character's gaze and posture track the ball.
We address the longstanding challenge of producing flexible, realistic humanoid character controllers that can perform diverse whole-body tasks involving object interactions. This challenge is central to a variety of fields, from graphics and animation to robotics and motor neuroscience. Our physics-based environment uses realistic actuation and first-person perception -including touch sensors and egocentric vision -with a view to producing active-sensing behaviors (e.g. gaze direction), transferability to real robots, and comparisons to the biology. We develop an integrated neuralnetwork based approach consisting of a motor primitive module, human demonstrations, and an instructed reinforcement learning regime with curricula and task variations. We demonstrate the utility of our approach for several tasks, including goal-conditioned box carrying and ball catching, and we characterize its behavioral robustness. The resulting controllers can be deployed in real-time on a standard PC. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35420d83-2f07-4ddd-a274-5f29976095a9Cited by top-tier papers47
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine et al.SIGGRAPH 2021 · 392 citations
- Stochastic Scene-Aware Motion PredictionMohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito et al.ICCV 2021 · 240 citations
- ASE: large-scale reusable adversarial skill embeddings for physically simulated charactersXue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine et al.SIGGRAPH 2022 · 217 citations
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersRuihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu et al.ICLR 2022 · 146 citations
Builds on1
Related papers
- Deep Compliant ControlSeunghwan Lee, Phil Sik Chang, Jehee LeeSIGGRAPH 2022 · 15 citations
- QuestEnvSim: Environment-Aware Simulated Motion Tracking from Sparse SensorsSunmin Lee, Sebastian Starke, Yuting Ye, Jungdam Won et al.SIGGRAPH 2023 · 31 citations
- Breathing Life Into Biomechanical User ModelsAleksi Ikkala, Florian Fischer, Markus Klar, Miroslav Bachinski et al.UIST 2022 · 33 citations
- Learning Soccer Juggling Skills with Layer-wise Mixture-of-ExpertsZhaoming Xie, Sebastian Starke, Hung Yu Ling, Michiel van de PanneSIGGRAPH 2022 · 13 citations
- DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelYonghao Zhang, Qiang He, Yanguang Wan, Yinda Zhang et al.AAAI 2025 · 10 citations
