Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion Prediction
Zhenyu Lou, Qiongjie Cui, Tuo Wang, Zhenbo Song, Luoming Zhang, Cheng Cheng, Haofan Wang, Xu Tang, Huaxia Li, Hong Zhou
摘要
Diverse human motion prediction (HMP) is a fundamental application in computer vision that has recently attracted considerable interest. Prior methods primarily focus on the stochastic nature of human motion, while neglecting the specific impact of the external environment, leading to the pronounced artifacts in prediction when applied to real-world scenarios. To fill this gap, this work introduces a novel task: predicting diverse human motion within real-world 3D scenes. In contrast to prior works, it requires harmonizing the deterministic constraints imposed by the surrounding 3D scenes with the stochastic aspect of human motion. For this purpose, we propose DiMoP3D , a diverse motion prediction framework with 3D scene awareness, which leverages the 3D point cloud and observed sequence to generate diverse and high-fidelity predictions. DiMoP3D can comprehend the 3D scene and determine the probable target objects and their desired interactive pose based on the historical motion. Then, it plans the obstacle-free trajectories toward these interested objects and generates diverse and physically consistent future motions. On top of that, DiMoP3D identifies deterministic factors in the scene and integrates them into stochastic modeling, making the diverse HMP in realistic scenes become a controllable stochastic
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative RefinementYude Zou, Junji Gong, Xing Gao, Zixuan Li 等ICLR 2026 · 被引用 3 次
- Scenemi: Motion In-Betweening for Modeling Human-Scene InteractionsInwoo Hwang, Bing Zhou, Young Min Kim, Jian Wang 等ICCV 2025
- FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent FrameworkLingzhou Mu, Qiang Wang, Fan Jiang, Mengchao Wang 等AAAI 2026
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai 等NeurIPS 2022 · 被引用 1,270 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
相关 Paper
- Motion Diversification NetworksHee Jae Kim, Eshed Ohn-BarCVPR 2024
- Towards Diverse and Natural Scene-aware 3D Human Motion SynthesisJingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan 等CVPR 2022 · 被引用 74 次
- Contextually Plausible and Diverse 3D Human Motion PredictionSadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson, Stephen Gould 等ICCV 2021 · 被引用 44 次
- Stochastic Multi-Person 3D Motion ForecastingSirui Xu, Yu-Xiong Wang, Liangyan GuiICLR 2023 · 被引用 3 次
- Contact-aware Human Motion ForecastingWei Mao, Miaomiao Liu, Richard I. Hartley, Mathieu SalzmannNeurIPS 2022 · 被引用 43 次
