Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation
Rawal Khirodkar, Visesh Chari, Amit Agrawal, Ambrish Tyagi
Abstract
A key assumption of top-down human pose estimation approaches is their expectation of having a single person/instance present in the input bounding box. This often leads to failures in crowded scenes with occlusions. We propose a novel solution to overcome the limitations of this fundamental assumption. Our Multi-Instance Pose Network (MIPNet) allows for predicting multiple 2D pose instances within a given bounding box. We introduce a Multi-Instance Modulation Block (MIMB) that can adaptively modulate channel-wise feature responses for each instance and is parameter efficient. We demonstrate the efficacy of our approach by evaluating on COCO, CrowdPose, and OCHuman datasets. Specifically, we achieve 70.0 AP on CrowdPose and 42.5 AP on OCHuman test sets, a significant improvement of 2.4 AP and 6.5 AP over the prior art, respectively. When using ground truth bounding boxes for inference, MIP-Net achieves an improvement of 0.7 AP on COCO, 0.9 AP on CrowdPose, and 9.1 AP on OCHuman validation sets compared to HRNet. Interestingly, when fewer, high confidence bounding boxes are used, HRNet’s performance degrades (by 5 AP) on OCHuman, whereas MIPNet maintains a relatively stable performance (drop of 1 AP) for the same inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c637c5c8-2f92-4197-b0e1-cc0134ee2717Cited by top-tier papers10
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- Rethinking pose estimation in crowds: overcoming the detection information bottleneck and ambiguityMu Zhou, Lucas Stoffl, Mackenzie Weygandt Mathis, Alexander MathisICCV 2023 · 28 citations
- Pose-Guided 3D Human Generation in Indoor SceneMinseok Kim, Changwoo Kang, Jeongin Park, Kyungdon JooAAAI 2023 · 5 citations
- Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionBavesh Balaji, Jerrin Bright, Yuhao Chen, Sirisha Rambhatla et al.NeurIPS 2024 · 4 citations
Builds on4
- You Only Train Once: Loss-Conditional Training of Deep NetworksAlexey Dosovitskiy, Josip DjolongaICLR 2020 · 96 citations
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi et al.CVPR 2020
- The Devil Is in the Details: Delving Into Unbiased Data Processing for Human Pose EstimationJunjie Huang, Zheng Zhu, Feng Guo, Guan HuangCVPR 2020
- Detection in Crowded Scenes: One Proposal, Multiple PredictionsXuangeng Chu, Anlin Zheng, Xiangyu Zhang, Jian SunCVPR 2020
Related papers
- Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous CircleMiroslav Purkrábek, Jiri MatasICCV 2025
- Semantic-aware Transfer with Instance-adaptive Parsing for Crowded Scenes Pose EstimationXuanhan Wang, Lianli Gao, Yan Dai, Yixuan Zhou et al.ACM MM 2021 · 14 citations
- AdaptivePose: Human Parts as Adaptive PointsYabo Xiao, Xiaojuan Wang, Dongdong Yu, Guoli Wang et al.AAAI 2022 · 25 citations
- InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose EstimationDahu Shi, Xing Wei, Xiaodong Yu, Wenming Tan et al.ACM MM 2021 · 40 citations
- Occluded Human Mesh RecoveryRawal Khirodkar, Shashank Tripathi, Kris KitaniCVPR 2022 · 74 citations
