SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs
Bi Xue, Hong Wu, Lei Chen, Chao Yang, Yiming Ma, Fei Ding, Zhen Wang, Liang Wang, Xiaoheng Mao, Ke Huang, Xialu Li, Peng Xia
Abstract
Serving deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from non-negligible costs and missing co-design opportunities. Such inefficiency makes them difficult to support complex model architectures, such as learned similarities and multi-task retrieval. In this paper, we present SilverTorch, a model-based serving system that brings all components into one unified model. It unifies model serving by replacing standalone indexing and filtering services with model layers. We propose a model-based GPU Bloom index for feature filtering and a fused Int8 ANN kernel for nearest neighbor search. Through co-design of the ANN search and feature filtering, we reduce GPU memory usage and eliminate computation. Benefiting from this design, we scale up retrieval by introducing an OverArch scoring layer and a multi-task retrieval with a Value Model to aggregate scores. These advancements improve the retrieval accuracy and enable future studies for serving more complex models. Our evaluation on industry-scale datasets shows that SilverTorch achieves up to 23.7× higher throughput compared to the state-of-the-art approaches. We also demonstrate that SilverTorch's solution is 13.35× more cost-efficient than CPU-based solution while improving accuracy via serving more complex models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46000006-1d14-46c5-a7bd-ef5e31b5d7cbBuilds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang et al.ICML 2024 · 200 citations
- Wukong: Towards a Scaling Law for Large-Scale RecommendationBuyun Zhang, Liang Luo, Yuxin Chen, Jade Nie et al.ICML 2024 · 108 citations
- CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUsHiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang et al.ICDE 2024 · 59 citations
- Retrieval with Learned SimilaritiesBailu Ding, Jiaqi ZhaiWWW 2025 · 3 citations
Related papers
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam et al.MICRO 2024 · 5 citations
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo et al.INFOCOM 2026
- Bagpipe: Accelerating Deep Recommendation Model TrainingSaurabh Agarwal, Chengpo Yan, Ziyi Zhang, Shivaram VenkataramanSOSP 2023 · 17 citations
- UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM ArchitectureSitian Chen, Amelie Chi Zhou, Yucheng Shi, Yusen Li et al.SC 2025 · 8 citations
- FlashANNS: GPU-Driven Asynchronous I/O Pipelining for Eliminating Storage-Compute Bottlenecks in Billion-Scale Similarity SearchYang Xiao, Mo Sun, Ziyu Song, Bing Tian et al.SIGMOD 2026 · 3 citations
