Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
Wenyi Liang, Jianchun Liu, Hongli Xu, Chunming Qiao, Liusheng Huang
摘要
Edge inference is a technology that enables real-time data processing and analysis on clients near the data source. To ensure compliance with the Service-Level Objectives (SLOs), such as a 30% latency reduction target, caching is usually adopted to reduce redundant computations in inference tasks on stream data. Due to task and data correlations, sharing cache information among clients can improve the inference performance. However, the non-independent and identically distributed (non-IID) nature of data across different clients and the long-tail distributions, where some classes have significantly more samples than others, will reduce cache hit ratios and increase latency. To address the aforementioned challenges, we propose an efficient inference framework, CoCa, which leverages a multi-client collaborative caching mechanism to accelerate edge inference. On the client side, the model is pre-set with multiple cache layers to achieve a quick inference. During inference, the model performs sequential lookups at cache layers activated by the edge server. On the server side, CoCa uses a two-dimensional global cache to periodically aggregate information from clients, mitigating the effects of non-IID data. For client cache allocation, CoCa first evaluates the importance of classes based on how frequently and recently their samples have been accessed. CoCa then selects frequently recurring classes to address long-tail distribution challenges. Finally, CoCa dynamically activates cache layers to balance lookup overhead and accuracy. Extensive experiments demonstrate that CoCa reduces inference latency by 23.0% to 45.2% on the VGG, ResNet and AST models with a slight loss of accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia 等MobiCom 2021 · 被引用 171 次
- EagleEye: wearable camera-based person identification in crowded urban spacesJuheon Yi, Sunghyun Choi, Youngki LeeMobiCom 2020 · 被引用 69 次
- VideoLT: Large-scale Long-tailed Video RecognitionXing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu 等ICCV 2021 · 被引用 51 次
- MergeSFL: Split Federated Learning with Feature Merging and Batch Size RegulationYunming Liao, Yang Xu, Hongli Xu, Lun Wang 等ICDE 2024 · 被引用 41 次
- Efficient Federated-Learning Model DebuggingAnran Li, Lan Zhang, Junhao Wang, Juntao Tan 等ICDE 2021 · 被引用 37 次
相关 Paper
- Intelligent Video Caching at Network Edge: A Multi-Agent Deep Reinforcement Learning ApproachFangxin Wang, Feng Wang, Jiangchuan Liu, Ryan Shea 等INFOCOM 2020 · 被引用 139 次
- ACBatch: Adaptive and Cooperative Batching for Edge InferenceZiming Yang, Zichuan Zheng, Liyou Deng, Shan Zhang 等INFOCOM 2025 · 被引用 2 次
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang 等INFOCOM 2025 · 被引用 3 次
- QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesKaihua Fu, Jiuchen Shi, Quan Chen, Ningxin Zheng 等SC 2022 · 被引用 7 次
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
