PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
Yunjae Lee, Hyeseong Kim, Minsoo Rhu
摘要
Training recommendation systems (RecSys) faces several challenges as it requires the “data preprocessing” stage to preprocess an ample amount of raw data and feed them to the GPU for training in a seamless manner. To sustain high training throughput, state-of-the-art solutions reserve a large fleet of CPU servers for preprocessing which incurs substantial deployment cost and power consumption. Our characterization reveals that prior CPU-centric preprocessing is bottlenecked on feature generation and feature normalization operations as it fails to reap out the abundant inter-/intra-feature parallelism in RecSys preprocessing. PreSto is a storage-centric preprocessing system leveraging In-Storage Processing (ISP), which offloads the bottlenecked preprocessing operations to our ISP units. We show that PreSto outperforms the baseline CPU-centric system with a 9.6× speedup in end-to-end preprocessing time, 4.3× enhancement in cost-efficiency, and 11.3× improvement in energy-efficiency on average for production-scale RecSys preprocessing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis 等HPCA 2025 · 被引用 6 次
- MinatoLoader: Accelerating Machine Learning Training Through Efficient Data PreprocessingRahma Nouaji, Stella Bitchebe, Ricardo Macedo, Oana BalmauEuroSys 2026 · 被引用 3 次
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu 等OSDI 2026
- Kirin: Efficient In-Storage Learned Compaction for LSM-Trees via System-Algorithm Co-DesignGuifeng Wang, Shengan Zheng, Penghao Sun, Jin Pu 等VLDB 2026
它引用的顶会 Paper28
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Facebook's Tectonic Filesystem: Efficiency from ExascaleSatadru Pan, Theano Stavrinos, Yunqiao Zhang, Atul Sikaria 等FAST 2021 · 被引用 110 次
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel 等ASPLOS 2021 · 被引用 100 次
相关 Paper
- RM-SSD: In-Storage Computing for Large-Scale Recommendation InferenceXuan Sun, Hu Wan, Qiao Li, Chia-Lin Yang 等HPCA 2022 · 被引用 33 次
- EagleRec: Edge-Scale Recommendation System Acceleration with Inter-Stage Parallelism Optimization on GPUsYongbo Yu, Fuxun Yu, Xiang Sheng, Chenchen Liu 等DAC 2023 · 被引用 1 次
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang 等SIGMOD 2025 · 被引用 8 次
- GLIST: Towards In-Storage Graph LearningCangyuan Li, Ying Wang, Cheng Liu, Shengwen Liang 等USENIX ATC 2021 · 被引用 53 次
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu 等VLDB 2021 · 被引用 85 次
