Proxy-Bridged Game Transformer for Interactive Extreme Motion Prediction
Yanwen Fang, Wenqi Jia, Xu Cao, Peng-Tao Jiang, Guodong Li, Jintai Chen
Abstract
Multi-person motion prediction becomes particularly challenging when handling highly interactive scenarios involving extreme motions. Previous works focused more on the case of 'moderate' motions (e.g., walking together), where predicting each pose in isolation often yields reasonable results. However, these approaches fall short in modeling extreme motions like lindy-hop dances, as they require a more comprehensive understanding of cross-person dependencies. To bridge this gap, we introduce Proxybridged Game Transformer (PGformer), a Transformerbased foundation model that captures the interactions driving extreme multi-person motions. PGformer incorporates a novel cross-query attention module to learn bidirectional dependencies between pose sequences and a proxy unit that subtly controls bidirectional spatial information flow. We evaluated PGformer on the challenging ExPI dataset, which involves large collaborative movements. Both quantitative and qualitative results demonstrate the superiority of PGformer in both short-and long-term predictions. We also test the proposed method on moderate movement datasets CMU-Mocap and MuPoTS-3D, generalizing PGformer to scenarios with more than two individuals with promising results. Code of PGformer is available at https://github.com/joyfang1106/pgformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionLingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 252 citations
- Structured Prediction Helps 3D Human Motion ModellingEmre Aksan, Manuel Kaufmann, Otmar HilligesICCV 2019 · 204 citations
- Space-Time-Separable Graph Convolutional Network for Pose ForecastingTheodoros Sofianos, Alessio Sampieri, Luca Franco, Fabio GalassoICCV 2021 · 188 citations
- Multi-Person 3D Motion Prediction with Multi-Range TransformersJiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong WangNeurIPS 2021 · 102 citations
Related papers
- Multi-Person Extreme Motion PredictionWen Guo, Xiaoyu Bie, Xavier Alameda-Pineda, Francesc Moreno-NoguerCVPR 2022 · 64 citations
- TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose EstimationJiajie Liu, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 27 citations
- Joint-Relation Transformer for Multi-Person Motion PredictionQingyao Xu, Weibo Mao, Jingze Gong, Chenxin Xu et al.ICCV 2023 · 24 citations
- TGFormer: Transformer with Track Query Group for Multi-Object TrackingRui Zeng, Yuanzhou Huang, Songwei PeiAAAI 2025 · 6 citations
- ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion GenerationLiang Xu, Ziyang Song, Dongliang Wang, Jing Su et al.ICCV 2023 · 100 citations
