VR-DANN: Real-Time Video Recognition via Decoder-Assisted Neural Network Acceleration
Zhuoran Song, Feiyang Wu, Xueyuan Liu, Jing Ke, Naifeng Jing, Xiaoyao Liang
Abstract
Nowadays, high-definition video object recognition (segmentation and detection) is not within the easy reach of a real-time task in a consumer SoC due to the limited on-chip computing power for neural network (NN) processing. Although many accelerators have been optimized heavily, they are still isolated from the intrinsic video compression expertise in a decoder. Given the fact that a great portion of frames can be dynamically reconstructed by a few key frames with high fidelity in a video, we envision that the recognition can also be reconstructed in a similar way so as to save a large amount of NN computing power. In this paper, we study the feasibility and efficiency of a novel decoder-assisted NN accelerator architecture for video recognition (VR-DANN) in a conventional SoC-styled design, which for the first time tightly couples the working principle of a video decoder with the NN accelerator to provide smooth high-definition video recognition experience. We leverage motion vectors, the simple tempo-spatial information already available in the decoding process to facilitate the recognition process, and propose a lightweight NN-based refinement scheme to suppress the non-pixel recognition noise. We also propose the corresponding microarchitecture design, which can be built upon any existing commercial IPs with minimal hardware overhead but significant speedup. Our experimental results show that the VR-DANN-parallel architecture achieves 2.9× performance improvement with less than 1% accuracy loss compared with the state-of-the-art "FAVOS" scheme widely used for video recognition. Compared with optical flow assisted "DFF" scheme, it can achieve 2.2× performance gain and 3% accuracy improvement. As to another "Euphrates" scheme, VR-DANN can achieve 40% performance gain and comparable accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e02840ee-79b0-4454-9e2e-87ef6b5fbd3eCited by top-tier papers5
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- Cicero: Addressing Algorithmic and Architectural Bottlenecks in Neural Rendering by Radiance Warping and Memory OptimizationsYu Feng, Zihan Liu, Jingwen Leng, Minyi Guo et al.ISCA 2024 · 18 citations
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting AccelerationChanghun Oh, Seongryong Oh, Jinwoo Hwang, Yoonsung Kim et al.ASPLOS 2026 · 6 citations
- Adaptive CHERI Compartmentalization for Heterogeneous AcceleratorsJianyi Cheng, A. Theodore Markettos, Alexandre Joannou, Paul Metzger et al.ISCA 2025 · 4 citations
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim et al.VLDB 2025 · 2 citations
Builds on2
Related papers
- E2SR: an end-to-end video CODEC assisted system for super resolution accelerationZhuoran Song, Zhongkai Yu, Naifeng Jing, Xiaoyao LiangDAC 2022 · 4 citations
- FastDVDnet: Towards Real-Time Deep Video Denoising Without Flow EstimationMatias Tassano, Julie Delon, Thomas VeitCVPR 2020
- PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge DevicesQihua Zhou, Song Guo, Jun Pan, Jiacheng Liang et al.AAAI 2023
- AccDecoder: Accelerated Decoding for Neural-enhanced Video AnalyticsTingting Yuan, Liang Mi, Weijun Wang, Haipeng Dai et al.INFOCOM 2023 · 25 citations
- Deep Learning Acceleration with Neuron-to-Memory TransformationMohsen Imani, Mohammad Samragh Razlighi, Yeseong Kim, Saransh Gupta et al.HPCA 2020 · 31 citations
