Forecasting of 3D Whole-Body Human Poses with Grasping Objects
Haitao Yan, Qiongjie Cui, Jiexin Xie, Shijie Guo
摘要
In the context of computer vision and human-robot interaction, forecasting 3D human poses is crucial for understanding human behavior and enhancing the predictive capabilities of intelligent systems. While existing methods have made significant progress, they often focus on predicting major body joints, overlooking fine-grained gestures and their interaction with objects. Human hand movements, particularly during object interactions, play a pivotal role and provide more precise expressions of human poses. This work fills this gap and introduces a novel paradigm: forecasting 3D whole-body human poses with a focus on grasping objects. This task involves predicting activities across all joints in the body and hands, encompassing the complexities of internal heterogeneity and external interactivity. To tackle these challenges, we also propose a novel approach: C<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>HOST, cross-context cross-modal consolidation for 3D whole-body pose forecasting, effectively handles the complexities of internal heterogeneity and external interactivity. C<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>HOST involves distinct steps, including the heterogeneous content encoding and alignment, and cross-modal feature learning and interaction. These enable us to predict activities across all body and hand joints, ensuring high-precision whole-body human pose prediction, even during object grasping. Extensive experiments on two benchmarks demonstrate that our model significantly enhances the accuracy of whole-body human motion prediction. The project page is available at https://sites.google.com/view/c3host.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- EgoAgent: A Joint Predictive Agent Model in Egocentric WorldsLu Chen, Yizhou Wang, Shixiang Tang, Qianhong Ma 等ICCV 2025 · 被引用 10 次
- FIction: 4D Future Interaction Prediction from VideoKumar Ashutosh, Georgios Pavlakos, Kristen GraumanCVPR 2025
- Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesTing Yu, Yi Lin, Jun Yu, Zhenyu Lou 等CVPR 2025
- CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangementYun Liu, Chengwen Zhang, Ruofan Xing, Bingda Tang 等CVPR 2025
它引用的顶会 Paper22
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 被引用 534 次
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 被引用 384 次
- MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionLingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 等ICCV 2021 · 被引用 252 次
- Human Motion Prediction via Spatio-Temporal InpaintingAlejandro Hernandez Ruiz, Jürgen Gall, Francesc MorenoICCV 2019 · 被引用 233 次
相关 Paper
- Expressive Forecasting of 3D Whole-Body Human MotionsPengxiang Ding, Qiongjie Cui, Haofan Wang, Min Zhang 等AAAI 2024 · 被引用 10 次
- DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelYonghao Zhang, Qiang He, Yanguang Wan, Yinda Zhang 等AAAI 2025 · 被引用 10 次
- Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands ModulatorGyeongsik MoonCVPR 2026 · 被引用 1 次
- COOP: Decoupling and Coupling of Whole-Body Grasping Pose GenerationYanzhao Zheng, Yunzhou Shi, Yuhao Cui, Zhongzhou Zhao 等ICCV 2023 · 被引用 8 次
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 被引用 201 次
