Pose-native Network Architecture Search for Multi-person Human Pose Estimation
Qian Bao, Wu Liu, Jun Hong, Lingyu Duan, Tao Mei
Abstract
Multi-person pose estimation has achieved great progress in recent years, even though, the precise prediction for occluded and invisible hard keypoints remains challenging. Most of the human pose estimation networks are equipped with an image classification-based pose encoder for feature extraction and a handcrafted pose decoder for high-resolution representations. However, the pose encoder might be sub-optimal because of the gap between image classification and pose estimation. The widely used multi-scale feature fusion in pose decoder is still coarse and cannot provide sufficient high-resolution details for hard keypoints. Neural Architecture Search (NAS) has shown great potential in many visual tasks to automatically search efficient networks. In this work, we present the Pose-native Network Architecture Search (PoseNAS) to simultaneously design a better pose encoder and pose decoder for pose estimation. Specifically, we directly search a data-oriented pose encoder with stacked searchable cells, which can provide an optimum feature extractor for the pose specific task. In the pose decoder, we exploit scale-adaptive fusion cells to promote rich information exchange across the multi-scale feature maps. Meanwhile, the pose decoder adopts a Fusion-and-Enhancement manner to progressively boost the high-resolution representations that are non-trivial for the precious prediction of hard keypoints. With the exquisitely designed search space and search strategy, PoseNAS can simultaneously search all modules in an end-to-end manner. PoseNAS achieves state-of-the-art performance on three public datasets, MPII, COCO, and PoseTrack, with small-scale parameters compared with the existing methods. Our best model obtains 76.7% mAP and 75.9% mAP on the COCO validation set and test set with only 33.6M parameters. Code and implementation are available at https://github.com/for-code0216/PoseNAS.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0a58d7fb-98f9-4c80-b4b1-8bb5ca87d4d0Cited by top-tier papers4
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 44 citations
- InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose EstimationDahu Shi, Xing Wei, Xiaodong Yu, Wenming Tan et al.ACM MM 2021 · 40 citations
- Neural Architecture Search for Joint Human Parsing and Pose EstimationDan Zeng, Yuhang Huang, Qian Bao, Junjie Zhang et al.ICCV 2021 · 20 citations
- Learning Latent Architectural Distribution in Differentiable Neural Architecture Search via Variational Information MaximizationYaoming Wang, Yuchen Liu, Wenrui Dai, Chenglin Li et al.ICCV 2021 · 9 citations
Related papers
- ViPNAS: Efficient Video Pose Estimation via Neural Architecture SearchLumin Xu, Yingda Guan, Sheng Jin, Wentao Liu et al.CVPR 2021
- HR-NAS: Searching Efficient High-Resolution Neural Architectures With Lightweight TransformersMingyu Ding, Xiaochen Lian, Linjie Yang, Peng Wang et al.CVPR 2021
- Auto-FPN: Automatic Network Architecture Adaptation for Object Detection Beyond ClassificationHang Xu, Lewei Yao, Zhenguo Li, Xiaodan Liang et al.ICCV 2019 · 197 citations
- Automatic Network Architecture Search for RGB-D Semantic SegmentationWenna Wang, Tao Zhuo, Xiuwei Zhang, Mingjun Sun et al.ACM MM 2023 · 6 citations
- BFBox: Searching Face-Appropriate Backbone and Feature Pyramid Network for Face DetectorYang Liu, Xu TangCVPR 2020
