HALP: Heuristic Aided Learned Preference Eviction Policy for YouTube Content Delivery Network
Zhenyu Song, Kevin Chen, Nuikhil Sarda, Deniz Altinbüken, Eugene Brevdo, Jimmy Coleman, Xiao Ju, Pawel Jurczyk, Richard Schooler, Ramki Gummadi
Abstract
Video streaming services are among the largest web applications in production, and a large source of downstream internet traffic. A large-scale video streaming service at Google, YouTube, leverages a Content Delivery Network (CDN) to serve its users. A key consideration in providing a seamless service is cache efficiency. In this work, we demonstrate machine learning techniques to improve the efficiency of YouTube's CDN DRAM cache. While many recently proposed learning-based caching algorithms show promising results, we identify and address three challenges blocking deployment of such techniques in a large-scale production environment: computation overhead for learning, robust byte miss ratio improvement, and measuring impact under production noise. We propose a novel caching algorithm, HALP, which achieves low CPU overhead and robust byte miss ratio improvement by augmenting a heuristic policy with machine learning. We also propose a production measurement method, impact distribution analysis, that can accurately measure the impact distribution of a new caching algorithm deployment in a noisy production environment.
HALP has been running in YouTube CDN production as a DRAM level eviction algorithm since early 2022 and has reliably reduced the byte miss during peak by an average of 9.1% while expending a modest CPU overhead of 1.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0df8bea6-6fd0-41d3-b1a9-a5119f40f335Cited by top-tier papers11
- Baleen: ML Admission & Prefetching for Flash CachesDaniel Lin-Kit Wong, Hao Wu, Carson Molder, Sathya Gunasekar et al.FAST 2024 · 26 citations
- 3L-Cache: Low Overhead and Precise Learning-based Eviction Policy for CachesWenbin Zhou, Zhixiong Niu, Yongqiang Xiong, Juan Fang et al.FAST 2025 · 16 citations
- Seer: Enabling Future-Aware Online Caching in Networked SystemsJason Lei, Vishal ShrivastavNSDI 2024 · 11 citations
- Learned Prefix Caching for Efficient LLM InferenceDongsheng Yang, Austin T. Li, Kai Li, Wyatt LloydNeurIPS 2025 · 8 citations
- Robustifying Learning-Augmented Caching Efficiently without Compromising 1-ConsistencyPeng Chen, Hailiang Zhao, Jiaji Zhang, Xueyan Tang et al.NeurIPS 2025 · 4 citations
Builds on4
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 193 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan et al.ICML 2020 · 108 citations
- Learning Cache Replacement with CACHEUSLiana V. Rodriguez, Farzana Beente Yusuf, Steven Lyons, Eysler Paz et al.FAST 2021 · 26 citations
Related papers
- RL-Bélády: A Unified Learning Framework for Content CachingGang Yan, Jian LiACM MM 2020 · 15 citations
- Intelligent Video Caching at Network Edge: A Multi-Agent Deep Reinforcement Learning ApproachFangxin Wang, Feng Wang, Jiangchuan Liu, Ryan Shea et al.INFOCOM 2020 · 139 citations
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 60 citations
- Darwin: Flexible Learning-based CDN CachingJiayi Chen, Nihal Sharma, Tarannum Khan, Shu Liu et al.SIGCOMM 2023 · 13 citations
- Robust Learning-Augmented Caching: An Experimental StudyJakub Chledowski, Adam Polak, Bartosz Szabucki, Konrad Tomasz ZolnaICML 2021 · 21 citations
