A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation
Qitao Zhao, Ce Zheng, Mengyuan Liu, Chen Chen
摘要
The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation, intractable computation and the non-causal problem. This can be attributed to their inherent inability to perceive spatial context as plain 2D joint coordinates carry no visual cues. To address this issue, we propose a straightforward yet powerful solution: leveraging the readily available intermediate visual representations produced by off-the-shelf (pre-trained) 2D pose detectors -no finetuning on the 3D task is even needed. The key observation is that, while the pose detector learns to localize 2D joints, such representations (e.g., feature maps) implicitly encode the joint-centric spatial context thanks to the regional operations in backbone networks. We design a simple baseline named Context-Aware PoseFormer to showcase its effectiveness. Without access to any temporal information, the proposed method significantly outperforms its context-agnostic counterpart, PoseFormer [77] , and other state-of-the-art methods using up to hundreds of video frames regarding both speed and precision. Project page: qitaozhao.github.io/ContextAware-PoseFormer * Work was done while Qitao was an intern mentored by Chen Chen. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear JacketChengxu Zuo, Yiming Wang, Lishuang Zhan, Shihui Guo 等CVPR 2024 · 被引用 19 次
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen 等ACM MM 2024 · 被引用 16 次
- Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and ModulationYinghao Wu, Chaoran Wang, Lu Yin, Shihui Guo 等NeurIPS 2024 · 被引用 11 次
- Toward Approaches to Scalability in 3D Human Pose EstimationJun-Hui Kim, Seong-Whan LeeNeurIPS 2024 · 被引用 5 次
- : Discrete Diffusion Model for Occluded 3D Human Pose EstimationWeiquan Wang, Jun Xiao, Chunping Wang, Wei Liu 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
相关 Paper
- PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose EstimationQitao Zhao, Ce Zheng, Mengyuan Liu, Pichao Wang 等CVPR 2023
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- ExtPose: Robust and Coherent Pose Estimation by Extending ViTsRongyu Chen, Li'an Zhuo, Linlin Yang, Qi Wang 等ICML 2025
- IVT: An End-to-End Instance-guided Video Transformer for 3D Pose EstimationZhongwei Qiu, Qiansheng Yang, Jian Wang, Dongmei FuACM MM 2022 · 被引用 7 次
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu 等ICCV 2023 · 被引用 322 次
