DeDelayed: Deleting Remote Inference Delay via On-Device Correction
Dan Jacobellis, Mateen Ulhaq, Fabien Racapé, Hyomin Choi, Neeraja J. Yadwadkar
摘要
Video comprises the vast majority of bits that are generated daily, and is the primary signal driving current innovations in robotics, remote sensing, and wearable technology. Yet, the most powerful video understanding models are too expensive for the resource-constrained platforms used in these applications. One approach is to offload inference to the cloud; this gives access to GPUs capable of processing high-resolution videos in real time. But even with reliable, high-bandwidth communication channels, the combined latency of video encoding, model inference, and round-trip communication prohibits use for certain real-time applications. The alternative is to use fully local inference; but this places extreme constraints on computational and power costs, requiring smaller models and lower resolution, leading to degraded accuracy. To address these challenges, we propose Dedelayed, a real-time inference system that divides computation between a remote model operating on delayed video frames and a local model with access to the current frame. The remote model is trained to make predictions on anticipated future frames, which the local model incorporates into its prediction for the current frame. The local and remote models are jointly optimized with an autoencoder that limits the transmission bitrate required by the available downlink communication channel. We evaluate Dedelayed on the task of real-time streaming video segmentation using the BDD100k driving dataset. For a round trip delay of 100 ms, Dedelayed improves performance by 6.4 mIoU compared to fully local inference and 9.8 mIoU compared to remote inference -- an equivalent improvement to using a model ten times larger. We release our training code, pretrained models, and python library at https://github.com/InterDigitalInc/dedelayed .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense PredictionHan Cai, Junyan Li, Muyan Hu, Chuang Gan 等ICCV 2023 · 被引用 265 次
- Real-Time Video Inference on Edge Devices via Adaptive Model StreamingMehrdad Khani Shirkoohi, Pouya Hamadanian, Arash Nasr-Esfahany, Mohammad AlizadehICCV 2021 · 被引用 58 次
相关 Paper
- Concerto: Client-server Orchestration for Real-Time Video AnalyticsChaoyang Li, Rui-Xiao Zhang, Tianchi Huang, Lianchen Jia 等ACM MM 2023 · 被引用 8 次
- Transtreaming: Adaptive Delay-aware Transformer for Real-time Streaming PerceptionXiang Zhang, Yufei Cui, Chenchen Fu, Zihao Wang 等AAAI 2025 · 被引用 2 次
- Server-Driven Video Streaming for Deep Learning InferenceKuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery 等SIGCOMM 2020 · 被引用 238 次
- AccDecoder: Accelerated Decoding for Neural-enhanced Video AnalyticsTingting Yuan, Liang Mi, Weijun Wang, Haipeng Dai 等INFOCOM 2023 · 被引用 25 次
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin 等CVPR 2020
