Task-Aware Encoder Control for Deep Video Compression
Xingtong Ge, Jixiang Luo, Xinjie Zhang, Tongda Xu, Guo Lu, Dailan He, Jing Geng, Yan Wang, Jun Zhang, Hongwei Qin
摘要
Prior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DVC decoder. Empirical evidence demonstrates that our method is applicable across multiple tasks with various existing pre-trained DVCs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bitrate for different tasks, with only one pre-trained decoder.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation CompressionZhimeng Huang, Rongao Yuan, Junlong Gao, Qi Mao 等CVPR 2026
- Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation CompressionSha Guo, Jing Chen, Zixuan Hu, Zhuo Chen 等ICLR 2025
- RL-RC-DoT: A Block-level RL agent for Task-Aware Video CompressionUri Gadot, Assaf Shocher, Shie Mannor, Gal Chechik 等CVPR 2025
它引用的顶会 Paper11
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode PredictionZhihao Hu, Guo Lu, Jinyang Guo, Shan Liu 等CVPR 2022 · 被引用 95 次
- YOLOV: Making Still Image Object Detectors Great at Video Object DetectionYuheng Shi, Naiyan Wang, Xiaojie GuoAAAI 2023 · 被引用 83 次
- Lossy Compression for Lossless PredictionYann Dubois, Benjamin Bloem-Reddy, Karen Ullrich, Chris J. MaddisonNeurIPS 2021 · 被引用 82 次
相关 Paper
- Complexity-guided Slimmable Decoder for Efficient Deep Video CompressionZhihao Hu, Dong XuCVPR 2023
- Offline and Online Optical Flow Enhancement for Deep Video CompressionChuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang 等AAAI 2024 · 被引用 35 次
- MMVC: Learned Multi-Mode Video Compression with Block-based Prediction Mode Selection and Density-Adaptive Entropy CodingBowen Liu, Yu Chen, Rakesh Chowdary Machineni, Shiyu Liu 等CVPR 2023
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 被引用 233 次
- DeepSVC: Deep Scalable Video Coding for Both Machine and Human VisionHongbin Lin, Bolin Chen, Zhichen Zhang, Jielian Lin 等ACM MM 2023 · 被引用 32 次
