RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
Uri Gadot, Assaf Shocher, Shie Mannor, Gal Chechik, Assaf Hallak
Abstract
Video encoders optimize compression for human perception by minimizing reconstruction error under bit-rate constraints. In many modern applications such as autonomous driving, an overwhelming majority of videos serve as input for AI systems performing tasks like object recognition or segmentation, rather than being watched by humans. It is therefore useful to optimize the encoder for a downstream task instead of for perceptual image quality. However, a major challenge is how to combine such downstream optimization with existing standard video encoders, which are highly efficient and popular. Here, we address this challenge by controlling the Quantization Parameters (QPs) at the macro-block level to optimize the downstream task. This granular control allows us to prioritize encoding for taskrelevant regions within each frame. We formulate this optimization problem as a Reinforcement Learning (RL) task, where the agent learns to balance long-term implications of choosing QPs on both task performance and bit-rate constraints. Notably, our policy does not require the downstream task as an input during inference, making it suitable for streaming applications and edge devices such as vehicles. We demonstrate significant improvements in two tasks, car detection, and ROI (saliency) encoding. Our approach improves task performance for a given bit rate compared to traditional task agnostic encoding methods, paving the way for more efficient task-aware video compression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7d63c90-9d4d-48a5-b772-77952a9a50bbCited by top-tier papers1
Ask how each one uses itBuilds on5
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- Task-aware Distributed Source Coding under Dynamic BandwidthPo-han Li, Sravan Kumar Ankireddy, Ruihan Philip Zhao, Hossein Nourkhiz Mahjoub et al.NeurIPS 2023 · 16 citations
- QS-NeRV: Real-Time Quality-Scalable Decoding with Neural Representation for VideosChang Wu, Guancheng Quan, Gang He, Xin-Quan Lai et al.ACM MM 2024 · 8 citations
- Task-Aware Encoder Control for Deep Video CompressionXingtong Ge, Jixiang Luo, Xinjie Zhang, Tongda Xu et al.CVPR 2024
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask LearningFisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian et al.CVPR 2020
Related papers
- SpotStream: Real-Time Video Transmission for Autonomous Driving via Small Object-Aware ROIZelin Song, Huanhuan Zhang, Pengcheng Zhang, Mingyue Zhao et al.INFOCOM 2026
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- Server-Driven Video Streaming for Deep Learning InferenceKuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery et al.SIGCOMM 2020 · 238 citations
- AMIS: Edge Computing Based Adaptive Mobile Video StreamingPhil K. Mu, Jinkai Zheng, Tom H. Luan, Lina Zhu et al.INFOCOM 2021 · 19 citations
- Learning-based Multi-Drone Network Edge Orchestration for Video AnalyticsChengyi Qu, Rounak Singh, Alicia Esquivel Morel, Prasad CalyamINFOCOM 2022 · 8 citations
