Joint-Relation Transformer for Multi-Person Motion Prediction
Qingyao Xu, Weibo Mao, Jingze Gong, Chenxin Xu, Siheng Chen, Weidi Xie, Ya Zhang, Yanfeng Wang
Abstract
Multi-person motion prediction is a challenging problem due to the dependency of motion on both individual past movements and interactions with other people. Transformer-based methods have shown promising results on this task, but they miss the explicit relation representation between joints, such as skeleton structure and pairwise distance, which is crucial for accurate interaction modeling. In this paper, we propose the Joint-Relation Transformer, which utilizes relation information to enhance interaction modeling and improve future motion prediction. Our relation information contains the relative distance and the intra-/inter-person physical constraints. To fuse relation and joint information, we design a novel joint-relation fusion layer with relation-aware attention to update both features. Additionally, we supervise the relation information by forecasting future distance. Experiments show that our method achieves a 13.4% improvement of 900ms VIM on 3DPW-SoMoF/RC and 17.8%/12.0% improvement of 3s MPJPE on CMU-Mpcap/MuPoTS-3D dataset. Code is available at https://github.com/MediaBrain-SJTU/JRTransformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0190d38f-3423-4409-b03a-454e073e9bf2Cited by top-tier papers9
- Language-Driven Interactive Traffic Trajectory GenerationJunkai Xia, Chenxin Xu, Qingyao Xu, Yanfeng Wang et al.NeurIPS 2024 · 27 citations
- Human Motion Prediction Under Unexpected PerturbationJiangbei Yue, Baiyi Li, Julien Pettré, Armin Seyfried et al.CVPR 2024 · 5 citations
- Ponimator: Unfolding Interactive Pose for Versatile Human-Human Interaction AnimationShaowei Liu, Chuan Guo, Bing Zhou, Jian WangICCV 2025 · 2 citations
- HUMOF: Human Motion Forecasting in Interactive Social ScenesCaiyi Sun, Yujing Sun, Xiao Han, Zemin Yang et al.ICLR 2026 · 2 citations
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan et al.CVPR 2024
Builds on18
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- Rethinking Graph Transformers with Spectral AttentionDevin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau et al.NeurIPS 2021 · 854 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
Related papers
- Multi-Person 3D Motion Prediction with Multi-Range TransformersJiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong WangNeurIPS 2021 · 102 citations
- Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal InteractionsYuanhong Zheng, Ruixuan Yu, Jian SunICCV 2025
- Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose EstimationHan Li, Bowen Shi, Wenrui Dai, Hongwei Zheng et al.AAAI 2023 · 76 citations
- Future Motion Dynamic Modeling via Hybrid Supervision for Multi-Person Motion Prediction Uncertainty ReductionYan Zhuang, Yanlu Cai, Weizhong Zhang, Cheng JinACM MM 2024 · 3 citations
- A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token CompletionJunkun Jiang, Jie Chen, Yike GuoACM MM 2022 · 8 citations
