Tracking People with 3D Representations
Jathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, Jitendra Malik
Abstract
We present a novel approach for tracking multiple people in video. Unlike past approaches which employ 2D representations, we focus on using 3D representations of people, located in three-dimensional space. To this end, we develop a method, Human Mesh and Appearance Recovery (HMAR) which in addition to extracting the 3D geometry of the person as a SMPL mesh, also extracts appearance as a texture map on the triangles of the mesh. This serves as a 3D representation for appearance that is robust to viewpoint and pose changes. Given a video clip, we first detect bounding boxes corresponding to people, and for each one, we extract 3D appearance, pose, and location information using HMAR. These embedding vectors are then sent to a transformer, which performs spatio-temporal aggregation of the representations over the duration of the sequence. The similarity of the resulting representations is used to solve for associations that assigns each person to a tracklet. We evaluate our approach on the Posetrack, MuPoTs and AVA datasets. We find that 3D representations are more effective than 2D representations for tracking in these settings, and we obtain state-of-the-art performance. Code and results are available at: https://brjathu.github.io/T3DP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3894b327-9f7d-4ca5-9e16-9cb2328c0786Cited by top-tier papers10
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- Tracking People by Predicting 3D Appearance, Location and PoseJathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, Jitendra MalikCVPR 2022 · 57 citations
- Human Mesh Recovery from Multiple ShotsGeorgios Pavlakos, Jitendra Malik, Angjoo KanazawaCVPR 2022 · 42 citations
- MobilePoser: Real-Time Full-Body Pose Estimation and 3D Human Translation from IMUs in Mobile Consumer DevicesVasco Xu, Chenfeng Gao, Henry Hoffmann, Karan AhujaUIST 2024 · 37 citations
- TEMPO: Efficient Multi-View Pose Estimation, Tracking, and ForecastingRohan Choudhury, Kris M. Kitani, László A. JeniICCV 2023 · 32 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
Related papers
- Animatable Virtual Humans: Learning Pose-Dependent Human Representations in UV Space for Interactive Performance SynthesisWieland Morgenstern, Milena T. Bagdasarian, Anna Hilsmann, Peter EisertIEEE VR 2024 · 7 citations
- Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from VideosHaoyuan Li, Haoye Dong, Hanchao Jia, Dong Huang et al.ICCV 2023 · 8 citations
- Shape-aware Multi-Person Pose Estimation from Multi-View ImagesZijian Dong, Jie Song, Xu Chen, Chen Guo et al.ICCV 2021 · 47 citations
- Visibility Aware Human-Object Interaction Tracking from Single RGB CameraXianghui Xie, Bharat Lal Bhatnagar, Gerard Pons-MollCVPR 2023
- PostureHMR: Posture Transformation for 3D Human Mesh RecoveryYu-Pei Song, Xiao Wu, Zhaoquan Yuanl, Jian-Jun Qiao et al.CVPR 2024 · 11 citations
