PetPS: Supporting Huge Embedding Models with Persistent Memory
Minhui Xie, Youyou Lu, Qing Wang, Yangyang Feng, Jiaqiang Liu, Kai Ren, Jiwu Shu
Abstract
Embedding models are effective for learning high-dimensional sparse data. Traditionally, they are deployed in DRAM parameter servers (PS) for online inference access. However, the ever-increasing model capacity makes this practice suffer from both high storage costs and long recovery time. Rapidly developing Persistent Memory (PM) offers new opportunities to PSs owing to its large capacity at low costs, as well as its persistence, while the application of PM also faces two challenges including high read latency and heavy CPU burden. To provide a low-cost but still high-performance parameter service for online inferences, we introduce PetPS, the first production-deployed PM parameter server. (1) To escape with high PM latency, PetPS introduces a PM hash index tailored for embedding model workloads, to minimize PM access. (2) To alleviate the CPU burden, PetPS offloads parameter gathering to NICs, to avoid CPU stalls when accessing parameters on PM and thus improve CPU efficiency. Our evaluation shows that PetPS can boost throughput by 1.3 -- 1.7X compared to PSs that use state-of-the-art PM hash indexes, or get 2.9 -- 5.5X latency reduction with the same throughput. Since 2020, PetPS has been deployed in Kuaishou, one world-leading short video company, and successfully reduced TCO by 30% without performance degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15fcdd81-d826-4d2a-9fe0-4b4905760353Cited by top-tier papers4
- Fast State Restoration in LLM Serving with HCacheShiwei Gao, Youmin Chen, Jiwu ShuEuroSys 2025 · 22 citations
- Revisiting Secondary Indexing in LSM-based Storage Systems with Persistent MemoryJing Wang, Youyou Lu, Qing Wang, Yuhao Zhang et al.USENIX ATC 2023 · 13 citations
- Sorting on Byte-Addressable Storage: The Resurgence of Tree StructureYing Zheng, Kian-Lee TanVLDB 2024 · 2 citations
- OMeGa: Boosting Large-scale Graph Embeddings with Heterogeneous Memory ProcessingPeng Fang, Siqiang Luo, Fang Wang, Bolong Zheng et al.ICDE 2025 · 1 citation
Builds on16
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- FlatStore: An Efficient Log-Structured Key-Value Storage Engine for Persistent MemoryYoumin Chen, Youyou Lu, Fan Yang, Qing Wang et al.ASPLOS 2020 · 166 citations
- Lock-free Concurrent Level Hashing for Persistent MemoryZhangyu Chen, Yu Hua, Bo Ding, Pengfei ZuoUSENIX ATC 2020 · 98 citations
- Viper: An Efficient Hybrid PMem-DRAM Key-Value StoreLawrence Benson, Hendrik Makait, Tilmann RablVLDB 2021 · 86 citations
- ROART: Range-query Optimized Persistent ARTShaonan Ma, Kang Chen, Shimin Chen, Mengxing Liu et al.FAST 2021 · 73 citations
Related papers
- SEPH: Scalable, Efficient, and Predictable Hashing on Persistent MemoryChao Wang, Junliang Hu, Tsun-Yu Yang, Yuhong Liang et al.OSDI 2023
- JPDHeap: A JVM Heap Design for PM-DRAM MemoriesLitong You, Tianxiao Gu, Shengan Zheng, Jianmei Guo et al.DAC 2021 · 2 citations
- MetoHash: A Memory-Efficient and Traffic-Optimized Hashing Index on Hybrid PMem-DRAM MemoriesZixiang Yu, Guangyang Deng, Zhirong Shen, Qiangsheng Su et al.SC 2025 · 1 citation
- APEX: A High-Performance Learned Index on Persistent MemoryBaotong Lu, Jialin Ding, Eric Lo, Umar Farooq Minhas et al.VLDB 2022 · 73 citations
- Plush: A Write-Optimized Persistent Log-Structured Hash-TableLukas Vogel, Alexander van Renen, Satoshi Imamura, Jana Giceva et al.VLDB 2022 · 27 citations
