Direct Multi-view Multi-person 3D Pose Estimation
Tao Wang, Jianfeng Zhang, Yujun Cai, Shuicheng Yan, Jiashi Feng
Abstract
We present Multi-view Pose transformer (MvP) for estimating multi-person 3D poses from multi-view images. Instead of estimating 3D joint locations from costly volumetric representation or reconstructing the per-person 3D pose from multiple detected 2D poses as in previous methods, MvP directly regresses the multi-person 3D poses in a clean and efficient way, without relying on intermediate tasks. Specifically, MvP represents skeleton joints as learnable query embeddings and let them progressively attend to and reason over the multi-view information from the input images to directly regress the actual 3D joint locations. To improve the accuracy of such a simple pipeline, MvP presents a hierarchical scheme to concisely represent query embeddings of multi-person skeleton joints and introduces an inputdependent query adaptation approach. Further, MvP designs a novel geometrically guided attention mechanism, called projective attention, to more precisely fuse the cross-view information for each joint. MvP also introduces a RayConv operation to integrate the view-dependent camera geometry into the feature representations for augmenting the projective attention. We show experimentally that our MvP model outperforms the state-of-the-art methods on several benchmarks while being much more efficient. Notably, it achieves 92.3% AP 25 on the challenging Panoptic dataset, improving upon the previous best approach [40] by 9.8%. MvP is general and also extendable to recovering human mesh represented by the SMPL model, thus useful for modeling multi-person body shapes. Code and models are available at https://github.com/sail-sg/mvp .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ab02229-6632-4abc-a182-af0bcbfeef00Cited by top-tier papers25
- Distribution-Aware Single-Stage Models for Multi-Person 3D Pose EstimationZitian Wang, Xuecheng Nie, Xiaochao Qu, Yunpeng Chen et al.CVPR 2022 · 44 citations
- PoseTriplet: Co-evolving 3D Human Pose Estimation, Imitation, and Hallucination under Self-supervisionKehong Gong, Bingbing Li, Jianfeng Zhang, Tao Wang et al.CVPR 2022 · 40 citations
- Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDARPeishan Cong, Yiteng Xu, Yiming Ren, Juze Zhang et al.AAAI 2023 · 37 citations
- Probabilistic Triangulation for Uncalibrated Multi-View 3D Human Pose EstimationBoyuan Jiang, Lei Hu, Shihong XiaICCV 2023 · 19 citations
- Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear JacketChengxu Zuo, Yiming Wang, Lishuang Zhan, Shihui Guo et al.CVPR 2024 · 19 citations
Builds on14
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin et al.ICLR 2020 · 379 citations
- Single-Stage Multi-Person Pose MachinesXuecheng Nie, Jiashi Feng, Jianfeng Zhang, Shuicheng YanICCV 2019 · 246 citations
- Cross View Fusion for 3D Human Pose EstimationHaibo Qiu, Chunyu Wang, Jingdong Wang, Naiyan Wang et al.ICCV 2019 · 242 citations
Related papers
- SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose EstimationVinkle Srivastav, Keqi Chen, Nicolas PadoyCVPR 2024 · 17 citations
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose EstimationAviral Chharia, Wenbo Gou, Haoye DongCVPR 2025
- Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose EstimationHan Li, Bowen Shi, Wenrui Dai, Hongwei Zheng et al.AAAI 2023 · 76 citations
- Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose EstimationKangkang Zhou, Lijun Zhang, Feng Lu, Xiang-Dong Zhou et al.ACM MM 2023 · 17 citations
- Shape-aware Multi-Person Pose Estimation from Multi-View ImagesZijian Dong, Jie Song, Xu Chen, Chen Guo et al.ICCV 2021 · 47 citations
