RM-SSD: In-Storage Computing for Large-Scale Recommendation Inference
Xuan Sun, Hu Wan, Qiao Li, Chia-Lin Yang, Tei-Wei Kuo, Chun Jason Xue
Abstract
To meet the strict service level agreement requirements of recommendation systems, the entire set of embeddings in recommendation systems needs to be loaded into the memory. However, as the model and dataset for production-scale recommendation systems scale up, the size of the embeddings is approaching the limit of memory capacity. Limited physical memory constrains the algorithms that can be trained and deployed, posing a severe challenge for deploying advanced recommendation systems. Recent studies offload the embedding lookups into SSDs, which targets the embedding-dominated recommendation models. This paper takes it one step further and proposes to offload the entire recommendation system into SSD with in-storage computing capability. The proposed SSD-side FPGA solution leverages a low-end FPGA to speed up both the embedding-dominated and MLP-dominated models with high resource efficiency. We evaluate the performance of the proposed solution with a prototype SSD. Results show that we can achieve 20-100× throughput improvement compared with the baseline SSD and 1.5-15× improvement compared with the state-of-art.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 265617d1-8a74-4e2f-96b6-0a9e4295da82Cited by top-tier papers12
- λ-IO: A Unified IO Stack for Computational StorageZhe Yang, Youyou Lu, Xiaojian Liao, Youmin Chen et al.FAST 2023 · 54 citations
- Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real SystemHongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park et al.HPCA 2024 · 26 citations
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin et al.MICRO 2024 · 26 citations
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 8 citations
- PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System InferencesPingyi Huo, Anusha Devulapally, Hasan Al Maruf, Minseo Park et al.MICRO 2024 · 6 citations
Related papers
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel et al.ASPLOS 2021 · 100 citations
- MaxEmbed: Maximizing SSD bandwidth utilization for huge embedding models servingRuwen Fan, Minhui Xie, Haodi Jiang, Youyou LuASPLOS 2024 · 1 citation
- MP-Rec: Hardware-Software Co-design to Enable Multi-path RecommendationSamuel Hsia, Udit Gupta, Bilge Acun, Newsha Ardalani et al.ASPLOS 2023 · 10 citations
- Fleche: an efficient GPU embedding cache for personalized recommendationsMinhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang et al.EuroSys 2022 · 24 citations
- Enabling Efficient Large Recommendation Model Training with Near CXL Memory ProcessingHaifeng Liu, Long Zheng, Yu Huang, Jingyi Zhou et al.ISCA 2024 · 24 citations
