ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search
Lumin Xu, Yingda Guan, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, Xiaogang Wang
摘要
Human pose estimation has achieved significant progress in recent years. However, most of the recent methods focus on improving accuracy using complicated models and ignoring real-time efficiency. To achieve a better trade-off between accuracy and efficiency, we propose a novel neural architecture search (NAS) method, termed ViP-NAS, to search networks in both spatial and temporal levels for fast online video pose estimation. In the spatial level, we carefully design the search space with five different dimensions including network depth, width, kernel size, group number, and attentions. In the temporal level, we search from a series of temporal feature fusions to optimize the total accuracy and speed across multiple video frames. To the best of our knowledge, we are the first to search for the temporal feature fusion and automatic computation allocation in videos. Extensive experiments demonstrate the effectiveness of our approach on the challenging COCO2017 and PoseTrack2018 datasets. Our discovered model family, S-ViPNAS and T-ViPNAS, achieve significantly higher inference speed (CPU real-time) without sacrificing the accuracy compared to the previous state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian 等CVPR 2022 · 被引用 132 次
- Pseudo-Labeled Auto-Curriculum Learning for Semi-Supervised Keypoint LocalizationCan Wang, Sheng Jin, Yingda Guan, Wentao Liu 等ICLR 2022 · 被引用 17 次
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 等AAAI 2025 · 被引用 8 次
- ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis UsersXiangyu Yin, Boyuan Yang, Weichen Liu, Qiyao Xue 等ICCV 2025 · 被引用 5 次
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper13
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 被引用 109 次
- Dynamic Kernel Distillation for Efficient Pose Estimation in VideosXuecheng Nie, Yuncheng Li, Linjie Luo, Ning Zhang 等ICCV 2019 · 被引用 76 次
相关 Paper
- Pose-native Network Architecture Search for Multi-person Human Pose EstimationQian Bao, Wu Liu, Jun Hong, Lingyu Duan 等ACM MM 2020 · 被引用 13 次
- AutoST: Efficient Neural Architecture Search for Spatio-Temporal PredictionTing Li, Junbo Zhang, Kainan Bao, Yuxuan Liang 等KDD 2020 · 被引用 86 次
- Searching for Two-Stream Models in Multivariate Space for Video RecognitionXinyu Gong, Heng Wang, Zheng Shou, Matt Feiszli 等ICCV 2021 · 被引用 9 次
- OPANAS: One-Shot Path Aggregation Network Architecture Search for Object DetectionTingting Liang, Yongtao Wang, Zhi Tang, Guosheng Hu 等CVPR 2021
- VONAS: Network Design in Visual Odometry using Neural Architecture SearchXing Cai, Lanqing Zhang, Chengyuan Li, Ge Li 等ACM MM 2020 · 被引用 8 次
