CoMotion: Concurrent Multi-person 3D Motion
Alejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter, Vladlen Koltun
Abstract
We introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and occlusions. Our model performs both strong per-frame detection and a learned pose update to track people from frame to frame. Rather than match detections across time, poses are updated directly from a new input image, which enables online tracking through occlusion. We train on numerous image and video datasets leveraging pseudolabeled annotations to produce a model that matches state-of-the-art systems in 3D pose estimation accuracy while being faster and more accurate in tracking multiple people through time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8322adb1-7518-4c4c-8dd1-7c6be3dceafeCited by top-tier papers7
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou et al.CVPR 2026 · 12 citations
- MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image AnimationXirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi et al.ICLR 2026 · 3 citations
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationxiang deng, Feng Gao, Yong Zhang, Youxin Pang et al.CVPR 2026 · 2 citations
- LAMP: Localization Aware Multi-camera People Tracking in Metric 3D WorldNan Yang, Julian Straub, Fan Zhang, Richard A. Newcombe et al.CVPR 2026 · 1 citation
- BRIDGE: Borderless Reconfiguration for Inclusive and Diverse Gameplay Experience via Embodiment TransformationHayato Saiki, Chunggi Lee, Hikari Takahashi, Tica Lin et al.CHI 2026 · 1 citation
Builds on26
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
Related papers
- Cross-View Tracking for Multi-Human 3D Pose Estimation at Over 100 FPSLong Chen, Haizhou Ai, Rui Chen, Zijie Zhuang et al.CVPR 2020
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- Learning Dynamics via Graph Neural Networks for Human Pose Estimation and TrackingYiding Yang, Zhou Ren, Haoxiang Li, Chunluan Zhou et al.CVPR 2021
- MultiPly: Reconstruction of Multiple People from Monocular Video in the WildZeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang et al.CVPR 2024
- Tracking People by Predicting 3D Appearance, Location and PoseJathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, Jitendra MalikCVPR 2022 · 57 citations
