Task-Aware Encoder Control for Deep Video Compression
Xingtong Ge, Jixiang Luo, Xinjie Zhang, Tongda Xu, Guo Lu, Dailan He, Jing Geng, Yan Wang, Jun Zhang, Hongwei Qin
Abstract
Prior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DVC decoder. Empirical evidence demonstrates that our method is applicable across multiple tasks with various existing pre-trained DVCs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bitrate for different tasks, with only one pre-trained decoder.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09586806-e228-45e2-95a9-ec6546c098b2Cited by top-tier papers3
- Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation CompressionZhimeng Huang, Rongao Yuan, Junlong Gao, Qi Mao et al.CVPR 2026
- Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation CompressionSha Guo, Jing Chen, Zixuan Hu, Zhuo Chen et al.ICLR 2025
- RL-RC-DoT: A Block-level RL agent for Task-Aware Video CompressionUri Gadot, Assaf Shocher, Shie Mannor, Gal Chechik et al.CVPR 2025
Builds on11
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode PredictionZhihao Hu, Guo Lu, Jinyang Guo, Shan Liu et al.CVPR 2022 · 95 citations
- YOLOV: Making Still Image Object Detectors Great at Video Object DetectionYuheng Shi, Naiyan Wang, Xiaojie GuoAAAI 2023 · 83 citations
- Lossy Compression for Lossless PredictionYann Dubois, Benjamin Bloem-Reddy, Karen Ullrich, Chris J. MaddisonNeurIPS 2021 · 82 citations
Related papers
- Complexity-guided Slimmable Decoder for Efficient Deep Video CompressionZhihao Hu, Dong XuCVPR 2023
- Offline and Online Optical Flow Enhancement for Deep Video CompressionChuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang et al.AAAI 2024 · 35 citations
- MMVC: Learned Multi-Mode Video Compression with Block-based Prediction Mode Selection and Density-Adaptive Entropy CodingBowen Liu, Yu Chen, Rakesh Chowdary Machineni, Shiyu Liu et al.CVPR 2023
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- DeepSVC: Deep Scalable Video Coding for Both Machine and Human VisionHongbin Lin, Bolin Chen, Zhichen Zhang, Jielian Lin et al.ACM MM 2023 · 32 citations
