ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search
Lumin Xu, Yingda Guan, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, Xiaogang Wang
Abstract
Human pose estimation has achieved significant progress in recent years. However, most of the recent methods focus on improving accuracy using complicated models and ignoring real-time efficiency. To achieve a better trade-off between accuracy and efficiency, we propose a novel neural architecture search (NAS) method, termed ViP-NAS, to search networks in both spatial and temporal levels for fast online video pose estimation. In the spatial level, we carefully design the search space with five different dimensions including network depth, width, kernel size, group number, and attentions. In the temporal level, we search from a series of temporal feature fusions to optimize the total accuracy and speed across multiple video frames. To the best of our knowledge, we are the first to search for the temporal feature fusion and automatic computation allocation in videos. Extensive experiments demonstrate the effectiveness of our approach on the challenging COCO2017 and PoseTrack2018 datasets. Our discovered model family, S-ViPNAS and T-ViPNAS, achieve significantly higher inference speed (CPU real-time) without sacrificing the accuracy compared to the previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49cd63af-201c-449b-b605-74ce5545f03aCited by top-tier papers9
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian et al.CVPR 2022 · 132 citations
- Pseudo-Labeled Auto-Curriculum Learning for Semi-Supervised Keypoint LocalizationCan Wang, Sheng Jin, Yingda Guan, Wentao Liu et al.ICLR 2022 · 17 citations
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu et al.AAAI 2025 · 8 citations
- ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis UsersXiangyu Yin, Boyuan Yang, Weichen Liu, Qiyao Xue et al.ICCV 2025 · 5 citations
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang et al.ACM MM 2024 · 2 citations
Builds on13
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 109 citations
- Dynamic Kernel Distillation for Efficient Pose Estimation in VideosXuecheng Nie, Yuncheng Li, Linjie Luo, Ning Zhang et al.ICCV 2019 · 76 citations
Related papers
- Pose-native Network Architecture Search for Multi-person Human Pose EstimationQian Bao, Wu Liu, Jun Hong, Lingyu Duan et al.ACM MM 2020 · 13 citations
- AutoST: Efficient Neural Architecture Search for Spatio-Temporal PredictionTing Li, Junbo Zhang, Kainan Bao, Yuxuan Liang et al.KDD 2020 · 86 citations
- Searching for Two-Stream Models in Multivariate Space for Video RecognitionXinyu Gong, Heng Wang, Zheng Shou, Matt Feiszli et al.ICCV 2021 · 9 citations
- OPANAS: One-Shot Path Aggregation Network Architecture Search for Object DetectionTingting Liang, Yongtao Wang, Zhi Tang, Guosheng Hu et al.CVPR 2021
- VONAS: Network Design in Visual Odometry using Neural Architecture SearchXing Cai, Lanqing Zhang, Chengyuan Li, Ge Li et al.ACM MM 2020 · 8 citations
