DeepSVC: Deep Scalable Video Coding for Both Machine and Human Vision
Hongbin Lin, Bolin Chen, Zhichen Zhang, Jielian Lin, Xu Wang, Tiesong Zhao
Abstract
Nowadays, end-to-end video coding for both machine and human vision has become an emerging research topic. In complicated systems such as large-scale internet of video things (IoVT), feature streams and video streams can be separately encoded and delivered for machine judgement and human viewing. In this paper, we propose a deep scalable video codec (DeepSVC) to support three-layer scalability from machine to human vision. First, we design a semantic layer that encodes semantic features extracted from the captured video for machine analysis. This layer employs a conditional semantic compression (CSC) method to remove redundancies between semantic features. Second, we design a structure layer that can be combined with semantic layer to predict the captured video at a low quality. This layer effectively estimates video frames based on semantic layer with an interlayer frame prediction (IFP) network. Third, we design a texture layer that can be combined with the above two layers to reconstruct high-quality video signals. This layer also takes advantage of the IFP network to improve its coding efficiency. In large-scale IoVT systems, DeepSVC can deliver semantic layer for regular use and transmit the other layers on demand. Experimental results indicate that the proposed DeepSVC outperforms popular codecs for machine and human vision. Compared with scalable extension of H.265/HEVC (SHVC), the proposed DeepSVC reduces average bit-per-pixel (bpp) by 25.51%/27.63%/59.87% at the same mAP/PSNR/MS-SSIM. Sourcecode is available at: https://github.com/LHB116/DeepSVC.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 40319129-aec6-4c64-b4c4-ebf1b2e35873Cited by top-tier papers4
- Task-Aware Encoder Control for Deep Video CompressionXingtong Ge, Jixiang Luo, Xinjie Zhang, Tongda Xu et al.CVPR 2024
- Firing Bits Where It Matters: Spiking-Guided Just Recognizable Distortion Modeling for Machine-Centric Video CodingWuyuan Xie, Zhenming Li, Yuwu Lu, Di Lin et al.AAAI 2026
- The Last Byte: Learning Just Enough for Machine-Oriented Image CompressionWuyuan Xie, Zhenming Li, Ye Liu, Jian Jin et al.AAAI 2026
- Neural-Centric Video Processing Pipeline for Unified Multi-Task InferenceSeyeon Lee, Juncheol Ye, Jaehong Kim, Dongsu HanCVPR 2026
Related papers
- Semantic Scalable Image Compression with Cross-Layer PriorsHanyue Tu, Li Li, Wengang Zhou, Houqiang LiACM MM 2021 · 16 citations
- Efficient Video Compression via Content-Adaptive Super-ResolutionMehrdad Khani Shirkoohi, Vibhaalakshmi Sivaraman, Mohammad AlizadehICCV 2021 · 68 citations
- Learning-Based Video Coding with Joint Deep Compression and EnhancementTiesong Zhao, Weize Feng, Hongji Zeng, Yiwen Xu et al.ACM MM 2022 · 24 citations
- Offline and Online Optical Flow Enhancement for Deep Video CompressionChuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang et al.AAAI 2024 · 35 citations
- Compressive sensing based asymmetric semantic image compression for resource-constrained IoT systemYujun Huang, Bin Chen, Jianghui Zhang, Han Qiu et al.DAC 2022 · 5 citations
