Dynamic Kernel Distillation for Efficient Pose Estimation in Videos
Xuecheng Nie, Yuncheng Li, Linjie Luo, Ning Zhang, Jiashi Feng
Abstract
Existing video-based human pose estimation methods extensively apply large networks onto every frame in the video to localize body joints, which suffer high computational cost and hardly meet the low-latency requirement in realistic applications. To address this issue, we propose a novel Dynamic Kernel Distillation (DKD) model to facilitate small networks for estimating human poses in videos, thus significantly lifting the efficiency. In particular, DKD introduces a light-weight distillator to online distill pose kernels via leveraging temporal cues from the previous frame in a one-shot feed-forward manner. Then, DKD simplifies body joint localization into a matching procedure between the pose kernels and the current frame, which can be efficiently computed via simple convolution. In this way, DKD fast transfers pose knowledge from one frame to provide compact guidance for body joint localization in the following frame, which enables utilization of small networks in video-based pose estimation. To facilitate the training process, DKD exploits a temporally adversarial training strategy that introduces a temporal discriminator to help generate temporally coherent pose kernels and pose estimation results within a long range. Experiments on Penn Action and Sub-JHMDB benchmarks demonstrate outperforming efficiency of DKD, specifically, 10× flops reduction and 2× speedup over previous best model, and its state-of-the-art accuracy. * This work was partly done while Xuecheng was an intern as Snap Inc. Small CNN Pose Kernel Distillator Matching Frame t-1 Frame t Small CNN Matching Frame t+1 Small CNN Pose Kernel Distillator Matching (a) Our DKD Model RNN or Optical Flow Large CNN Classification Frame t-1 RNN or Optical Flow Large CNN Frame t Large CNN Frame t+1 Classification Classification (b) The Traditional Model
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31c9e453-468d-4517-901f-ec21c6878dc9Cited by top-tier papers19
- Online Knowledge Distillation for Efficient Pose EstimationZheng Li, Jingwen Ye, Mingli Song, Ying Huang et al.ICCV 2021 · 123 citations
- Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionMiao Zhang, Jie Liu, Yifei Wang, Yongri Piao et al.ICCV 2021 · 112 citations
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu et al.CVPR 2022 · 76 citations
- Salient-to-Broad Transition for Video Person Re-identificationShutao Bai, Bingpeng Ma, Hong Chang, Rui Huang et al.CVPR 2022 · 71 citations
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 33 citations
Related papers
- DistilPose: Tokenized Pose Regression with Heatmap DistillationSuhang Ye, Yingyi Zhang, Jie Hu, Liujuan Cao et al.CVPR 2023
- Adaptive Decoupled Pose Knowledge DistillationJie Xu, Shanshan Zhang, Jian YangACM MM 2023 · 1 citation
- Knowledge Distillation for 6D Pose Estimation by Aligning Distributions of Local PredictionsShuxuan Guo, Yinlin Hu, José M. Álvarez, Mathieu SalzmannCVPR 2023
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin et al.CVPR 2020
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu et al.CVPR 2026 · 24 citations
