Accelerating Deep Learning Classification with Error-controlled Approximate-key Caching
Alessandro Finamore, James Roberts, Massimo Gallo, Dario Rossi
摘要
While Deep Learning (DL) technologies are a promising tool to solve networking problems that map to classification tasks, their computational complexity is still too high with respect to real-time traffic measurements requirements. To reduce the DL inference cost, we propose a novel caching paradigm, that we named approximate-key caching, which returns approximate results for lookups of selected input based on cached DL inference results. While approximate cache hits alleviate DL inference workload and increase the system throughput, they however introduce an approximation error. As such, we couple approximate-key caching with an error-correction principled algorithm, that we named auto-refresh. We analytically model our caching system performance for classic LRU and ideal caches, we perform a trace-driven evaluation of the expected performance, and we compare the benefits of our proposed approach with the state-of-the-art similarity caching – this testifies the practical interest of our proposal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 被引用 99 次
- Similarity Caching: Theory and AlgorithmsMichele Garetto, Emilio Leonardi, Giovanni NegliaINFOCOM 2020 · 被引用 32 次
- GRADES: Gradient Descent for Similarity CachingAnirudh Sabnis, Tareq Si Salem, Giovanni Neglia, Michele Garetto 等INFOCOM 2021 · 被引用 14 次
相关 Paper
- Autonomous Unknown-Application Filtering and Labeling for DL-based Traffic Classifier UpdateJielun Zhang, Fuhao Li, Feng Ye, Hongyu WuINFOCOM 2020 · 被引用 120 次
- ISAC: In-Switch Approximate Cache for IoT Object Detection and RecognitionWenquan Xu, Zijian Zhang, Haoyu Song, Shuxin Liu 等INFOCOM 2023 · 被引用 7 次
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 被引用 193 次
- LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupXiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang 等MobiCom 2023 · 被引用 29 次
- Caravan: Practical Online Learning of In-Network ML Models with Labeling AgentsQizheng Zhang, Ali Imran, Enkeleda Bardhi, Tushar Swamy 等OSDI 2024 · 被引用 18 次
