PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge Devices
Qihua Zhou, Song Guo, Jun Pan, Jiacheng Liang, Zhenda Xu, Jingren Zhou
摘要
Real-time video perception tasks are often challenging over the resource-constrained edge devices due to the concerns of accuracy drop and hardware overhead, where saving computations is the key to performance improvement. Existing methods either rely on domain-specific neural chips or priorly searched models, which require specialized optimization according to different task properties. In this work, we propose a general and task-independent Patch Automatic Skip Scheme (PASS), a novel end-to-end learning pipeline to support diverse video perception settings by decoupling acceleration and tasks. The gist is to capture the temporal similarity across video frames and skip the redundant computations at patch level, where the patch is a non-overlapping square block in visual. PASS equips each convolution layer with a learnable gate to selectively determine which patches could be safely skipped without degrading model accuracy. As to each layer, a desired gate needs to make flexible skip decisions based on intermediate features without any annotations, which cannot be achieved by conventional supervised learning paradigm. To address this challenge, we are the first to construct a tough self-supervisory procedure for optimizing these gates, which learns to extract contrastive representation, i.e., distinguishing similarity and difference, from frame sequence. These high-capacity gates can serve as a plug-and-play module for convolutional neural network (CNN) backbones to implement patch-skippable architectures, and automatically generate proper skip strategy to accelerate different video-based downstream tasks, e.g., outperforming the state-of-the-art MobileHumanPose (MHP) in 3D pose estimation and FairMOT in multiple object tracking, by up to 9.43 times and 12.19 times speedups, respectively. By directly processing the raw data of frames, PASS can generalize to real-time video streams on commodity edge devices, e.g., NVIDIA Jetson Nano, with efficient performance in realistic deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 被引用 368 次
- HarDNet: A Low Memory Traffic NetworkPing Chao, Chao-Yang Kao, Yu-Shan Ruan, Chien-Hsiang Huang 等ICCV 2019 · 被引用 303 次
- A variegated look at 5G in the wild: performance, power, and QoE implicationsArvind Narayanan, Xumiao Zhang, Ruiyang Zhu, Ahmad Hassan 等SIGCOMM 2021 · 被引用 259 次
相关 Paper
- BEVSA: A Real-Time Bird's-Eye-View Semantic Segmentation Accelerator for Multi-Camera SystemSangho Lee, Jueun Jung, Wuyoung Jang, Jihyeon Hwang 等DAC 2025
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin 等CVPR 2022 · 被引用 31 次
- VR-DANN: Real-Time Video Recognition via Decoder-Assisted Neural Network AccelerationZhuoran Song, Feiyang Wu, Xueyuan Liu, Jing Ke 等MICRO 2020 · 被引用 29 次
- SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame MaskZekun Ai, Xiaotong Luo, Yanyun Qu, Yuan XieACM MM 2024 · 被引用 2 次
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
