Multi-Person 3D Motion Prediction with Multi-Range Transformers
Jiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong Wang
Abstract
We propose a novel framework for multi-person 3D motion trajectory prediction. Our key observation is that a human's action and behaviors may highly depend on the other persons around. Thus, instead of predicting each human pose trajectory in isolation, we introduce a Multi-Range Transformers model which contains of a local-range encoder for individual motion and a global-range encoder for social interactions. The Transformer decoder then performs prediction for each person by taking a corresponding pose as a query which attends to both local and global-range encoder features. Our model not only outperforms state-of-the-art methods on long-term 3D motion prediction, but also generates diverse social interactions. More interestingly, our model can even predict 15-person motion simultaneously by automatically dividing the persons into different interaction groups. Project page with code is available at https://jiashunwang.github.io/MRT/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20beb25a-6fa6-4808-8705-e5644efcef9dCited by top-tier papers25
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
- Social Diffusion: Long-term Multiple Human Motion AnticipationJulian Tanke, Linguang Zhang, Amy Zhao, Chengcheng Tang et al.ICCV 2023 · 39 citations
- Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingWentao Bao, Lele Chen, Libing Zeng, Zhong Li et al.ICCV 2023 · 34 citations
- InterControl: Zero-shot Human Interaction Generation by Controlling Every JointZhenzhi Wang, Jingbo Wang, Yixuan Li, Dahua Lin et al.NeurIPS 2024 · 27 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
Related papers
- Stochastic Multi-Person 3D Motion ForecastingSirui Xu, Yu-Xiong Wang, Liangyan GuiICLR 2023 · 3 citations
- Joint-Relation Transformer for Multi-Person Motion PredictionQingyao Xu, Weibo Mao, Jingze Gong, Chenxin Xu et al.ICCV 2023 · 24 citations
- Future Motion Dynamic Modeling via Hybrid Supervision for Multi-Person Motion Prediction Uncertainty ReductionYan Zhuang, Yanlu Cai, Weizhong Zhang, Cheng JinACM MM 2024 · 3 citations
- PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersZhongwei Qiu, Qiansheng Yang, Jian Wang, Haocheng Feng et al.CVPR 2023
- Trajectory Unified Transformer for Pedestrian Trajectory PredictionLiushuai Shi, Le Wang, Sanping Zhou, Gang HuaICCV 2023 · 100 citations
