Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation
Rawal Khirodkar, Visesh Chari, Amit Agrawal, Ambrish Tyagi
摘要
A key assumption of top-down human pose estimation approaches is their expectation of having a single person/instance present in the input bounding box. This often leads to failures in crowded scenes with occlusions. We propose a novel solution to overcome the limitations of this fundamental assumption. Our Multi-Instance Pose Network (MIPNet) allows for predicting multiple 2D pose instances within a given bounding box. We introduce a Multi-Instance Modulation Block (MIMB) that can adaptively modulate channel-wise feature responses for each instance and is parameter efficient. We demonstrate the efficacy of our approach by evaluating on COCO, CrowdPose, and OCHuman datasets. Specifically, we achieve 70.0 AP on CrowdPose and 42.5 AP on OCHuman test sets, a significant improvement of 2.4 AP and 6.5 AP over the prior art, respectively. When using ground truth bounding boxes for inference, MIP-Net achieves an improvement of 0.7 AP on COCO, 0.9 AP on CrowdPose, and 9.1 AP on OCHuman validation sets compared to HRNet. Interestingly, when fewer, high confidence bounding boxes are used, HRNet’s performance degrades (by 5 AP) on OCHuman, whereas MIPNet maintains a relatively stable performance (drop of 1 AP) for the same inputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
- Rethinking pose estimation in crowds: overcoming the detection information bottleneck and ambiguityMu Zhou, Lucas Stoffl, Mackenzie Weygandt Mathis, Alexander MathisICCV 2023 · 被引用 28 次
- Pose-Guided 3D Human Generation in Indoor SceneMinseok Kim, Changwoo Kang, Jeongin Park, Kyungdon JooAAAI 2023 · 被引用 5 次
- Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionBavesh Balaji, Jerrin Bright, Yuhao Chen, Sirisha Rambhatla 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper4
- You Only Train Once: Loss-Conditional Training of Deep NetworksAlexey Dosovitskiy, Josip DjolongaICLR 2020 · 被引用 96 次
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi 等CVPR 2020
- The Devil Is in the Details: Delving Into Unbiased Data Processing for Human Pose EstimationJunjie Huang, Zheng Zhu, Feng Guo, Guan HuangCVPR 2020
- Detection in Crowded Scenes: One Proposal, Multiple PredictionsXuangeng Chu, Anlin Zheng, Xiangyu Zhang, Jian SunCVPR 2020
相关 Paper
- Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous CircleMiroslav Purkrábek, Jiri MatasICCV 2025
- Semantic-aware Transfer with Instance-adaptive Parsing for Crowded Scenes Pose EstimationXuanhan Wang, Lianli Gao, Yan Dai, Yixuan Zhou 等ACM MM 2021 · 被引用 14 次
- AdaptivePose: Human Parts as Adaptive PointsYabo Xiao, Xiaojuan Wang, Dongdong Yu, Guoli Wang 等AAAI 2022 · 被引用 25 次
- InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose EstimationDahu Shi, Xing Wei, Xiaodong Yu, Wenming Tan 等ACM MM 2021 · 被引用 40 次
- Occluded Human Mesh RecoveryRawal Khirodkar, Shashank Tripathi, Kris KitaniCVPR 2022 · 被引用 74 次
