Accelerating Deep Learning Classification with Error-controlled Approximate-key Caching
Alessandro Finamore, James Roberts, Massimo Gallo, Dario Rossi
Abstract
While Deep Learning (DL) technologies are a promising tool to solve networking problems that map to classification tasks, their computational complexity is still too high with respect to real-time traffic measurements requirements. To reduce the DL inference cost, we propose a novel caching paradigm, that we named approximate-key caching, which returns approximate results for lookups of selected input based on cached DL inference results. While approximate cache hits alleviate DL inference workload and increase the system throughput, they however introduce an approximation error. As such, we couple approximate-key caching with an error-correction principled algorithm, that we named auto-refresh. We analytically model our caching system performance for classic LRU and ideal caches, we perform a trace-driven evaluation of the expected performance, and we compare the benefits of our proposed approach with the state-of-the-art similarity caching – this testifies the practical interest of our proposal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7010bd1-36cc-4fa9-95da-0fa6df7e42b0Cited by top-tier papers1
Ask how each one uses itBuilds on5
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis et al.MobiCom 2020 · 312 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Return of the Lernaean Hydra: Experimental Evaluation of Data Series Approximate Similarity SearchKarima Echihabi, Kostas Zoumpatianos, Themis Palpanas, Houda BenbrahimVLDB 2020 · 99 citations
- Similarity Caching: Theory and AlgorithmsMichele Garetto, Emilio Leonardi, Giovanni NegliaINFOCOM 2020 · 32 citations
- GRADES: Gradient Descent for Similarity CachingAnirudh Sabnis, Tareq Si Salem, Giovanni Neglia, Michele Garetto et al.INFOCOM 2021 · 14 citations
Related papers
- Autonomous Unknown-Application Filtering and Labeling for DL-based Traffic Classifier UpdateJielun Zhang, Fuhao Li, Feng Ye, Hongyu WuINFOCOM 2020 · 120 citations
- ISAC: In-Switch Approximate Cache for IoT Object Detection and RecognitionWenquan Xu, Zijian Zhang, Haoyu Song, Shuxin Liu et al.INFOCOM 2023 · 7 citations
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 193 citations
- LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupXiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang et al.MobiCom 2023 · 29 citations
- Caravan: Practical Online Learning of In-Network ML Models with Labeling AgentsQizheng Zhang, Ali Imran, Enkeleda Bardhi, Tushar Swamy et al.OSDI 2024 · 18 citations
