Future Motion Dynamic Modeling via Hybrid Supervision for Multi-Person Motion Prediction Uncertainty Reduction
Yan Zhuang, Yanlu Cai, Weizhong Zhang, Cheng Jin
Abstract
Multi-person motion prediction remains a challenging problem due to the intricate motion dynamics and complex interpersonal interactions, where uncertainty escalates rapidly across the forecasting horizon. Existing approaches always overlook the motion dynamic modeling among the prediction frames to reduce the uncertainty, but leave it entirely up to the deep neural networks, which lacks a dynamic inductive bias, leading to suboptimal performance. This paper addresses this limitation by proposing an effective multi-person motion prediction method named Hybrid Supervision Transformer (HSFormer), which formulates the dynamic modeling within the prediction horizon as a novel hybrid supervision task. To be precise, our method performs a rolling predicting process equipped with a hybrid supervision mechanism, which enforces the model to be able to predict the pose in the next frames based on the (typically error-contained) earlier predictions. Addition to the standard supervision loss, two self and auxiliary supervision mechanisms, which minimize the distance of the predictions with error-contained inputs and the predictions with error-free inputs (ground truth) and guide the model to make accurate predictions based on the ground truth, are introduced to improve the robustness of our model to the input deviation in inference and stabilize the training process, respectively. The optimization techniques, such as stop-gradient, are extended to our model to improve the training efficiency. Furthermore, we develop a fine-grained spatio-temporal correlation capture module to assist the feature learning and reduce the uncertainties arising from the intricate and varying interactions among the individuals. Our approach achieves state-of-the-art results on multiple multi-person datasets in both short- and long-term prediction.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ef2becad-0365-4b5b-ba01-3dcc6915480cCited by top-tier papers1
Ask how each one uses itRelated papers
- Multi-Person 3D Motion Prediction with Multi-Range TransformersJiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong WangNeurIPS 2021 · 102 citations
- Joint-Relation Transformer for Multi-Person Motion PredictionQingyao Xu, Weibo Mao, Jingze Gong, Chenxin Xu et al.ICCV 2023 · 24 citations
- Human MotionFormer: Transferring Human Motions with Vision TransformersHongyu Liu, Xintong Han, Chenbin Jin, Lihui Qian et al.ICLR 2023 · 5 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- Proxy-Bridged Game Transformer for Interactive Extreme Motion PredictionYanwen Fang, Wenqi Jia, Xu Cao, Peng-Tao Jiang et al.ICCV 2025
