Multimodal Sense-Informed Forecasting of 3D Human Motions
Zhenyu Lou, Qiongjie Cui, Haofan Wang, Xu Tang, Hong Zhou
Abstract
Predicting future human pose is a fundamental application for machine intelligence, which drives robots to plan their behavior and paths ahead of time to seamlessly accomplish human-robot collaboration in real-world 3D scenarios. Despite encouraging results, existing approaches rarely consider the effects of the external scene on the motion sequence, leading to pronounced artifacts and physical implausibilities in the predictions. To address this limitation, this work introduces a novel multi-modal sense-informed motion prediction approach, which conditions high-fidelity generation on two modal information: external 3D scene, and internal human gaze, and is able to recognize their salience for future human activity. Furthermore, the gaze information is regarded as the human intention, and combined with both motion and scene features, we construct a ternary intention-aware attention to supervise the generation to match where the human wants to reach. Meanwhile, we introduce semantic coherence-aware attention to explicitly distinguish the salient point clouds and the underlying ones, to ensure a reasonable interaction of the generated sequence with the 3D scene. On two real-world benchmarks, the proposed method achieves state-of-the-art performance both in 3D human pose and trajectory prediction. More detailed results are available on the page: https://sites.google.com/view/cvpr2024sif3d.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 249b3045-6b3b-4c6b-bcbe-77dbea7470eeCited by top-tier papers5
- Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion PredictionZhenyu Lou, Qiongjie Cui, Tuo Wang, Zhenbo Song et al.NeurIPS 2024 · 10 citations
- Progressive Guessing to Fixed Point: Rethinking Human Motion Prediction with Deep Equilibrium ModelsDong Wei, Huaijiang Sun, Fan Liu, Yuhui ZhengCVPR 2026 · 1 citation
- FIction: 4D Future Interaction Prediction from VideoKumar Ashutosh, Georgios Pavlakos, Kristen GraumanCVPR 2025
- Shape My Moves: Text-Driven Shape-Aware Synthesis of Human MotionsTing-Hsuan Liao, Yi Zhou, Yu Shen, Chun-Hao Paul Huang et al.CVPR 2025
- EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingSongpengcheng Xia, Yu Zhang, Zhuo Su, Xiaozheng Zheng et al.CVPR 2025
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
Related papers
- Contact-aware Human Motion ForecastingWei Mao, Miaomiao Liu, Richard I. Hartley, Mathieu SalzmannNeurIPS 2022 · 43 citations
- Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed RealityTaewook Ha, Woojin Cho, Dooyoung Kim, Woontack WooIEEE VR 2026
- Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesTing Yu, Yi Lin, Jun Yu, Zhenyu Lou et al.CVPR 2025
- Multi-Agent Long-Term 3D Human Pose Forecasting via Interaction-Aware Trajectory ConditioningJaewoo Jeong, Daehee Park, Kuk-Jin YoonCVPR 2024
- Synthesizing Long-Term 3D Human Motion and Interaction in 3D ScenesJiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu et al.CVPR 2021
