MaxEmbed: Maximizing SSD bandwidth utilization for huge embedding models serving
Ruwen Fan, Minhui Xie, Haodi Jiang, Youyou Lu
Abstract
Deep learning recommendation models (DLRMs) have gained widespread application across search, advertising, and ecommerce. Still, DLRMs present notable challenges as they depend heavily on large embedding tables to represent sparse features in recommendation systems. This raises concerns about both memory capacity and cost. Solid-state drives (SSDs) offer a cost-effective solution with a significantly larger capacity, but they introduce read amplification issues because of the mismatch between embedding size and SSD read granularity. Prior SSD embedding storage systems aim to tackle these challenges by employing hypergraph partitioning to co-locate co-appearing embeddings onto the same SSD page, alleviating read amplification. However, this approach has a drawback as it divides embeddings into completely disjoint clusters, limiting potential combinations between embeddings.
In response to this limitation, we introduce MaxEmbed. Capitalizing on the extensive storage capacity of SSDs, Max-Embed effectively mines relationships between storage combinations of embeddings with replication, thereby enhancing the effective bandwidth of SSDs. Additionally, MaxEmbed incorporates a corresponding online service module for embedding query request handling, leveraging two key optimizations to reduce the overhead brought by replication. Our evaluations demonstrate that MaxEmbed boosts SSD embedding serving throughput by up to 18.7% under various settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 418a03bb-9240-4f85-a876-83812a1ae493Cited by top-tier papers2
- Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation LinkingTuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao et al.ASPLOS 2025 · 1 citation
- Discard-Based Garbage Collection for Distributed Log-Structured Storage Systems in ByteDanceRunhua Bian, Liqiang Zhang, Jinxin Liu, Jiacheng Zhang et al.FAST 2026
Builds on8
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang et al.AAAI 2021 · 131 citations
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel et al.ASPLOS 2021 · 100 citations
- MERCI: efficient embedding reduction on commodity hardware via sub-query memoizationYejin Lee, Seong Hoon Seo, Hyunji Choi, Hyoung Uk Sul et al.ASPLOS 2021 · 34 citations
- RM-SSD: In-Storage Computing for Large-Scale Recommendation InferenceXuan Sun, Hu Wan, Qiao Li, Chia-Lin Yang et al.HPCA 2022 · 33 citations
Related papers
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis et al.ASPLOS 2022 · 65 citations
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 2 citations
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang et al.AAAI 2026 · 1 citation
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li et al.DAC 2024 · 9 citations
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai et al.OSDI 2023 · 23 citations
