Forecasting of 3D Whole-Body Human Poses with Grasping Objects
Haitao Yan, Qiongjie Cui, Jiexin Xie, Shijie Guo
Abstract
In the context of computer vision and human-robot interaction, forecasting 3D human poses is crucial for understanding human behavior and enhancing the predictive capabilities of intelligent systems. While existing methods have made significant progress, they often focus on predicting major body joints, overlooking fine-grained gestures and their interaction with objects. Human hand movements, particularly during object interactions, play a pivotal role and provide more precise expressions of human poses. This work fills this gap and introduces a novel paradigm: forecasting 3D whole-body human poses with a focus on grasping objects. This task involves predicting activities across all joints in the body and hands, encompassing the complexities of internal heterogeneity and external interactivity. To tackle these challenges, we also propose a novel approach: C<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>HOST, cross-context cross-modal consolidation for 3D whole-body pose forecasting, effectively handles the complexities of internal heterogeneity and external interactivity. C<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>HOST involves distinct steps, including the heterogeneous content encoding and alignment, and cross-modal feature learning and interaction. These enable us to predict activities across all body and hand joints, ensuring high-precision whole-body human pose prediction, even during object grasping. Extensive experiments on two benchmarks demonstrate that our model significantly enhances the accuracy of whole-body human motion prediction. The project page is available at https://sites.google.com/view/c3host.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EgoAgent: A Joint Predictive Agent Model in Egocentric WorldsLu Chen, Yizhou Wang, Shixiang Tang, Qianhong Ma et al.ICCV 2025 · 10 citations
- FIction: 4D Future Interaction Prediction from VideoKumar Ashutosh, Georgios Pavlakos, Kristen GraumanCVPR 2025
- Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesTing Yu, Yi Lin, Jun Yu, Zhenyu Lou et al.CVPR 2025
- CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangementYun Liu, Chengwen Zhang, Ruofan Xing, Bingda Tang et al.CVPR 2025
Builds on22
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionLingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 252 citations
- Human Motion Prediction via Spatio-Temporal InpaintingAlejandro Hernandez Ruiz, Jürgen Gall, Francesc MorenoICCV 2019 · 233 citations
Related papers
- Expressive Forecasting of 3D Whole-Body Human MotionsPengxiang Ding, Qiongjie Cui, Haofan Wang, Min Zhang et al.AAAI 2024 · 10 citations
- DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelYonghao Zhang, Qiang He, Yanguang Wan, Yinda Zhang et al.AAAI 2025 · 10 citations
- Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands ModulatorGyeongsik MoonCVPR 2026 · 1 citation
- COOP: Decoupling and Coupling of Whole-Body Grasping Pose GenerationYanzhao Zheng, Yunzhou Shi, Yuhao Cui, Zhongzhou Zhao et al.ICCV 2023 · 8 citations
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
