IVT: An End-to-End Instance-guided Video Transformer for 3D Pose Estimation
Zhongwei Qiu, Qiansheng Yang, Jian Wang, Dongmei Fu
摘要
Video 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal information from sequential 2D poses, which cannot model the contextual depth feature effectively since the visual depth features are lost in the step of 2D pose estimation. In this paper, we simplify the paradigm into an end-to-end framework, Instance-guided Video Transformer (IVT), which enables learning spatiotemporal contextual depth information from visual features effectively and predicts 3D poses directly from video frames. In particular, we firstly formulate video frames as a series of instance-guided tokens and each token is in charge of predicting the 3D pose of a human instance. These tokens contain body structure information since they are extracted by the guidance of joint offsets from the human center to the corresponding body joints. Then, these tokens are sent into IVT for learning spatiotemporal contextual depth. In addition, we propose a cross-scale instance-guided attention mechanism to handle the variational scales among multiple persons. Finally, the 3D poses of each person are decoded from instance-guided tokens by coordinate regression. Experiments on three widely-used 3D pose estimation benchmarks show that the proposed IVT achieves state-of-the-art performances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose EstimationQitao Zhao, Ce Zheng, Mengyuan Liu, Chen ChenNeurIPS 2023 · 被引用 40 次
- Pedestrian-Centric 3D Pre-collision Pose and Shape Estimation from Dashcam PerspectiveMeijun Wang, Yu Meng, Zhongwei Qiu, Chao Zheng 等NeurIPS 2024 · 被引用 1 次
- PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersZhongwei Qiu, Qiansheng Yang, Jian Wang, Haocheng Feng 等CVPR 2023
它引用的顶会 Paper18
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai 等ICCV 2019 · 被引用 504 次
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 被引用 368 次
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 被引用 360 次
相关 Paper
- ExtPose: Robust and Coherent Pose Estimation by Extending ViTsRongyu Chen, Li'an Zhuo, Linlin Yang, Qi Wang 等ICML 2025
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose EstimationWanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 被引用 4 次
- Optimizing Human Pose Estimation Through Focused Human and Joint RegionsYingying Jiao, Zhigang Wang, Zhenguang Liu, Shaojing Fan 等AAAI 2025 · 被引用 4 次
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 被引用 2 次
