MaxEmbed: Maximizing SSD bandwidth utilization for huge embedding models serving
Ruwen Fan, Minhui Xie, Haodi Jiang, Youyou Lu
摘要
Deep learning recommendation models (DLRMs) have gained widespread application across search, advertising, and ecommerce. Still, DLRMs present notable challenges as they depend heavily on large embedding tables to represent sparse features in recommendation systems. This raises concerns about both memory capacity and cost. Solid-state drives (SSDs) offer a cost-effective solution with a significantly larger capacity, but they introduce read amplification issues because of the mismatch between embedding size and SSD read granularity. Prior SSD embedding storage systems aim to tackle these challenges by employing hypergraph partitioning to co-locate co-appearing embeddings onto the same SSD page, alleviating read amplification. However, this approach has a drawback as it divides embeddings into completely disjoint clusters, limiting potential combinations between embeddings.
In response to this limitation, we introduce MaxEmbed. Capitalizing on the extensive storage capacity of SSDs, Max-Embed effectively mines relationships between storage combinations of embeddings with replication, thereby enhancing the effective bandwidth of SSDs. Additionally, MaxEmbed incorporates a corresponding online service module for embedding query request handling, leveraging two key optimizations to reduce the overhead brought by replication. Our evaluations demonstrate that MaxEmbed boosts SSD embedding serving throughput by up to 18.7% under various settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation LinkingTuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao 等ASPLOS 2025 · 被引用 1 次
- Discard-Based Garbage Collection for Distributed Log-Structured Storage Systems in ByteDanceRunhua Bian, Liqiang Zhang, Jinxin Liu, Jiacheng Zhang 等FAST 2026
它引用的顶会 Paper8
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang 等AAAI 2021 · 被引用 131 次
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel 等ASPLOS 2021 · 被引用 100 次
- MERCI: efficient embedding reduction on commodity hardware via sub-query memoizationYejin Lee, Seong Hoon Seo, Hyunji Choi, Hyoung Uk Sul 等ASPLOS 2021 · 被引用 34 次
- RM-SSD: In-Storage Computing for Large-Scale Recommendation InferenceXuan Sun, Hu Wan, Qiao Li, Chia-Lin Yang 等HPCA 2022 · 被引用 33 次
相关 Paper
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis 等ASPLOS 2022 · 被引用 65 次
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 被引用 2 次
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang 等AAAI 2026 · 被引用 1 次
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li 等DAC 2024 · 被引用 9 次
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai 等OSDI 2023 · 被引用 23 次
