PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
Yunjae Lee, Hyeseong Kim, Minsoo Rhu
Abstract
Training recommendation systems (RecSys) faces several challenges as it requires the “data preprocessing” stage to preprocess an ample amount of raw data and feed them to the GPU for training in a seamless manner. To sustain high training throughput, state-of-the-art solutions reserve a large fleet of CPU servers for preprocessing which incurs substantial deployment cost and power consumption. Our characterization reveals that prior CPU-centric preprocessing is bottlenecked on feature generation and feature normalization operations as it fails to reap out the abundant inter-/intra-feature parallelism in RecSys preprocessing. PreSto is a storage-centric preprocessing system leveraging In-Storage Processing (ISP), which offloads the bottlenecked preprocessing operations to our ISP units. We show that PreSto outperforms the baseline CPU-centric system with a 9.6× speedup in end-to-end preprocessing time, 4.3× enhancement in cost-efficiency, and 11.3× improvement in energy-efficiency on average for production-scale RecSys preprocessing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis et al.HPCA 2025 · 6 citations
- MinatoLoader: Accelerating Machine Learning Training Through Efficient Data PreprocessingRahma Nouaji, Stella Bitchebe, Ricardo Macedo, Oana BalmauEuroSys 2026 · 3 citations
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu et al.OSDI 2026
- Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-DesignGuifeng Wang, Shengan Zheng, Penghao Sun, Jin Pu et al.VLDB 2026
Builds on28
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
- Facebook's Tectonic Filesystem: Efficiency from ExascaleSatadru Pan, Theano Stavrinos, Yunqiao Zhang, Atul Sikaria et al.FAST 2021 · 110 citations
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel et al.ASPLOS 2021 · 100 citations
Related papers
- RM-SSD: In-Storage Computing for Large-Scale Recommendation InferenceXuan Sun, Hu Wan, Qiao Li, Chia-Lin Yang et al.HPCA 2022 · 33 citations
- EagleRec: Edge-Scale Recommendation System Acceleration with Inter-Stage Parallelism Optimization on GPUsYongbo Yu, Fuxun Yu, Xiang Sheng, Chenchen Liu et al.DAC 2023 · 1 citation
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang et al.SIGMOD 2025 · 8 citations
- GLIST: Towards In-Storage Graph LearningCangyuan Li, Ying Wang, Cheng Liu, Shengwen Liang et al.USENIX ATC 2021 · 53 citations
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu et al.VLDB 2021 · 85 citations
