Kangaroo: Caching Billions of Tiny Objects on Flash
Sara McAllister, Benjamin Berg, Julian Tutuncu-Macias, Juncheng Yang, Sathya Gunasekar, Jimmy Lu, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger
Abstract
Many social-media and IoT services have very large working sets consisting of billions of tiny (≈100 B) objects. Large, flash-based caches are important to serving these working sets at acceptable monetary cost. However, caching tiny objects on flash is challenging for two reasons: (i) SSDs can read/write data only in multi-KB "pages" that are much larger than a single object, stressing the limited number of times flash can be written; and (ii) very few bits per cached object can be kept in DRAM without losing flash's cost advantage. Unfortunately, existing flash-cache designs fall short of addressing these challenges: write-optimized designs require too much DRAM, and DRAM-optimized designs require too many flash writes.
We present Kangaroo, a new flash-cache design that optimizes both DRAM usage and flash writes to maximize cache performance while minimizing cost. Kangaroo combines a large, set-associative cache with a small, log-structured cache. The set-associative cache requires minimal DRAM, while the log-structured cache minimizes Kangaroo's flash writes. Experiments using traces from Facebook and Twitter show that Kangaroo achieves DRAM usage close to the best prior DRAM-optimized design, flash writes close to the best prior write-optimized design, and miss ratios better than both. Kangaroo's design is Pareto-optimal across a range of allowed write rates, DRAM sizes, and flash sizes, reducing misses by 29% over the state of the art. These results are corroborated with a test deployment of Kangaroo in a production flash cache at Facebook.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d6e62b7-94e8-428f-8ce9-8d1e14139f5fCited by top-tier papers35
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang et al.ASPLOS 2022 · 103 citations
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 60 citations
- FIFO queues are all you need for cache evictionJuncheng Yang, Yazhuo Zhang, Ziyue Qiu, Yao Yue et al.SOSP 2023 · 54 citations
- Designing Cloud Servers for Lower CarbonJaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene et al.ISCA 2024 · 49 citations
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An et al.OSDI 2026 · 40 citations
Builds on5
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 193 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Segcache: a memory-efficient and scalable in-memory key-value cache for small objectsJuncheng Yang, Yao Yue, Rashmi VinayakNSDI 2021 · 70 citations
Related papers
- Nemo: A Low-Write-Amplification Cache for Tiny Objects on Log-Structured Flash DevicesXufeng Yang, Tingting Tan, Jingxin Hu, Congming Gao et al.ASPLOS 2026
- Improving Performance of Flash Based Key-Value Stores Using Storage Class Memory as a Volatile Memory ExtensionHiwot Tadese Kassa, Jason Akers, Mrinmoy Ghosh, Zhichao Cao et al.USENIX ATC 2021 · 46 citations
- All-Flash Array Key-Value Cache for Large ObjectsJinhyung Koo, Jinwook Bae, Minjeong Yuk, Seonggyun Oh et al.EuroSys 2023 · 9 citations
- MIDAS: Minimizing Write Amplification in Log-Structured Systems through Adaptive Group Number and Size ConfigurationSeonggyun Oh, Jeeyun Kim, Soyoung Han, Jaeho Kim et al.FAST 2024 · 16 citations
- TSCache: An Efficient Flash-based Caching Scheme for Time-series Data WorkloadsJian Liu, Kefei Wang, Feng ChenVLDB 2021 · 12 citations
