TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting
Rohan Choudhury, Kris M. Kitani, László A. Jeni
Abstract
Existing volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a robust spatiotemporal representation, improving pose accuracy while also tracking and forecasting human pose. We significantly reduce computation compared to the state-of-the-art by recurrently computing per-person 2D pose features, fusing both spatial and temporal information into a single representation. In doing so, our model is able to use spatiotemporal context to predict more accurate human poses without sacrificing efficiency. We further use this representation to track human poses over time as well as predict future poses. Finally, we demonstrate that our model is able to generalize across datasets without scene-specific fine-tuning. TEMPO achieves 10% better MPJPE with a 33× improvement in FPS compared to TesseTrack on the challenging CMU Panoptic Studio dataset. Our code and demos are available at https://rccchoudhury.github.io/tempo2023/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86a540b2-8b27-4a99-921e-d7e411537729Cited by top-tier papers4
- SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose EstimationVinkle Srivastav, Keqi Chen, Nicolas PadoyCVPR 2024 · 17 citations
- FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial SensorsRuonan Zheng, Jiawei Fang, Yuan Yao, Xiaoxia Gao et al.CHI 2025 · 5 citations
- DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose EstimationTony Danjun Wang, Tolga Birdal, Nassir Navab, Lennart BastianICML 2026
- Multi-Agent Long-Term 3D Human Pose Forecasting via Interaction-Aware Trajectory ConditioningJaewoo Jeong, Daehee Park, Kuk-Jin YoonCVPR 2024
Builds on19
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- Monocular, One-stage, Regression of Multiple 3D PeopleYu Sun, Qian Bao, Wu Liu, Yili Fu et al.ICCV 2021 · 335 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
Related papers
- TesseTrack: End-to-End Learnable Multi-Person Articulated 3D Pose TrackingN. Dinesh Reddy, Laurent Guigues, Leonid Pishchulin, Jayan Eledath et al.CVPR 2021
- Distribution-Aware Single-Stage Models for Multi-Person 3D Pose EstimationZitian Wang, Xuecheng Nie, Xiaochao Qu, Yunpeng Chen et al.CVPR 2022 · 44 citations
- Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal InteractionsYuanhong Zheng, Ruixuan Yu, Jian SunICCV 2025
- CoMotion: Concurrent Multi-person 3D MotionAlejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter et al.ICLR 2025
- Cross-View Tracking for Multi-Human 3D Pose Estimation at Over 100 FPSLong Chen, Haizhou Ai, Rui Chen, Zijie Zhuang et al.CVPR 2020
