SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs
Bi Xue, Hong Wu, Lei Chen, Chao Yang, Yiming Ma, Fei Ding, Zhen Wang, Liang Wang, Xiaoheng Mao, Ke Huang, Xialu Li, Peng Xia
摘要
Serving deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from non-negligible costs and missing co-design opportunities. Such inefficiency makes them difficult to support complex model architectures, such as learned similarities and multi-task retrieval. In this paper, we present SilverTorch, a model-based serving system that brings all components into one unified model. It unifies model serving by replacing standalone indexing and filtering services with model layers. We propose a model-based GPU Bloom index for feature filtering and a fused Int8 ANN kernel for nearest neighbor search. Through co-design of the ANN search and feature filtering, we reduce GPU memory usage and eliminate computation. Benefiting from this design, we scale up retrieval by introducing an OverArch scoring layer and a multi-task retrieval with a Value Model to aggregate scores. These advancements improve the retrieval accuracy and enable future studies for serving more complex models. Our evaluation on industry-scale datasets shows that SilverTorch achieves up to 23.7× higher throughput compared to the state-of-the-art approaches. We also demonstrate that SilverTorch's solution is 13.35× more cost-efficient than CPU-based solution while improving accuracy via serving more complex models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang 等ICML 2024 · 被引用 200 次
- Wukong: Towards a Scaling Law for Large-Scale RecommendationBuyun Zhang, Liang Luo, Yuxin Chen, Jade Nie 等ICML 2024 · 被引用 108 次
- CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUsHiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang 等ICDE 2024 · 被引用 59 次
- Retrieval with Learned SimilaritiesBailu Ding, Jiaqi ZhaiWWW 2025 · 被引用 3 次
相关 Paper
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam 等MICRO 2024 · 被引用 5 次
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo 等INFOCOM 2026
- Bagpipe: Accelerating Deep Recommendation Model TrainingSaurabh Agarwal, Chengpo Yan, Ziyi Zhang, Shivaram VenkataramanSOSP 2023 · 被引用 17 次
- UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM ArchitectureSitian Chen, Amelie Chi Zhou, Yucheng Shi, Yusen Li 等SC 2025 · 被引用 8 次
- FlashANNS: GPU-Driven Asynchronous I/O Pipelining for Eliminating Storage-Compute Bottlenecks in Billion-Scale Similarity SearchYang Xiao, Mo Sun, Ziyu Song, Bing Tian 等SIGMOD 2026 · 被引用 3 次
