SpotStream: Real-Time Video Transmission for Autonomous Driving via Small Object-Aware ROI
Zelin Song, Huanhuan Zhang, Pengcheng Zhang, Mingyue Zhao, Congkai An, Anfu Zhou, Liang Liu
Abstract
Cloud-based autonomous driving relies on low-latency video transmission for accurate downstream perception tasks. However, low-latency video streaming solutions, optimized for human Quality of Experience (QoE), are misaligned with the needs of machine vision tasks. While Region-of-Interest (ROI) encoding aims to bridge this gap, existing methods suffer from two critical limitations: first, insufficient protection for small and vulnerable objects, which are often missed by detection mechanisms or given uniform, inadequate resource allocation; second, prohibitive encoding overhead from fine-grained partitioning, leading to poor system robustness in dynamic network environments. To address these issues, we propose SpotStream, a novel low-latency video transmission framework. SpotStream features a dual-stream importance prediction network to ensure comprehensive coverage and prioritized protection for small objects. It couples this with an efficient, multi-level ROI assignment strategy that significantly reduces encoding overhead by simplifying region structure while focusing resources on critical targets. Comprehensive evaluations demonstrate that SpotStream substantially outperforms state-of-the-art baselines. Notably, under challenging real-world network traces, Spot-Stream improves the F1 score for critical small objects by an average of 12.9% while maintaining excellent system robustness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 21eebd2e-982c-44f4-bb6d-b3cd8564c489Related papers
- Saliency-Guided Foveated Video Encoding for Low-Latency and Immersive Cloud VRZe Wu, Ahmad Alhilal, Yuk Hang Tsui, Wen Jye Chai et al.IEEE VR 2026
- RL-RC-DoT: A Block-level RL agent for Task-Aware Video CompressionUri Gadot, Assaf Shocher, Shie Mannor, Gal Chechik et al.CVPR 2025
- Frame Complexity-Aware Foveated Video Encoding for Real-time High-Quality StreamingZe Wu, Ahmad Alhilal, Yuk Hang Tsui, Matti Siekkinen et al.IEEE VR 2026
- FovRL: Joint Foveation and Quality Control for Immersive VR Streaming Using Reinforcement LearningYuk Hang Tsui, Ze Wu, Ahmad Alhilal, Matti Siekkinen et al.WWW 2026 · 1 citation
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 45 citations
