FAERY: An FPGA-accelerated Embedding-based Retrieval System
Chaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han, Yuhang Jiang, Ding Tang, Zilong Wang, Kai Chen, Chuanxiong Guo
摘要
Embedding-based retrieval (EBR) is widely used in recommendation systems to retrieve thousands of relevant candidates from a large corpus with millions or more items. A good EBR system needs to achieve both high throughput and low latency, as high throughput usually means cost saving and low latency improves user experience. Unfortunately, the performance of existing CPU-and GPU-based EBR are far from optimal due to their inherent architectural limitations.
In this paper, we first study how an ideal yet practical EBR system works, and then design FAERY, an FPGA-accelerated EBR, which achieves the optimal performance of the practically ideal EBR system. FAERY is composed of three key components: It uses a high bandwidth HBM for memory bandwidth-intensive corpus scanning, a data parallelism approach for similarity calculation, and a pipeline-based approach for K-selection. To further reduce hardware resources, FAERY introduces a filter to early drop the non-Top-K items. Experiments show that the degraded FAERY with the same memory bandwidth of GPU still achieves 1.21×-12.27× lower latency and up to 4.29× higher throughput under a latency target of 10 ms than GPU-based EBR.
- This work is done while Chaoliang Zeng, Ding Tang, and Zilong Wang are interns in ByteDance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
- Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered PipeliningZili Zhang, Fangyue Liu, Gang Huang, Xuanzhe Liu 等NSDI 2024 · 被引用 37 次
- FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated LearningJunxue Zhang, Xiaodian Cheng, Wei Wang, Liu Yang 等NSDI 2023 · 被引用 29 次
- DF-GAS: a Distributed FPGA-as-a-Service Architecture towards Billion-Scale Graph-based Approximate Nearest Neighbor SearchShulin Zeng, Zhenhua Zhu, Jun Liu, Haoyu Zhang 等MICRO 2023 · 被引用 24 次
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai 等OSDI 2023 · 被引用 23 次
它引用的顶会 Paper5
- Degree-Quant: Quantization-Aware Training for Graph Neural NetworksShyam Anil Tailor, Javier Fernández-Marqués, Nicholas Donald LaneICLR 2021 · 被引用 180 次
- SONG: Approximate Nearest Neighbor Search on GPUWeijie Zhao, Shulong Tan, Ping LiICDE 2020 · 被引用 103 次
- FPGA-Accelerated Compactions for LSM-based Key-Value StoreTeng Zhang, Jianying Wang, Xuntao Cheng, Hao Xu 等FAST 2020 · 被引用 99 次
- Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load BalancingChaoliang Zeng, Layong Luo, Teng Zhang, Zilong Wang 等NSDI 2022 · 被引用 97 次
- QuickNN: Memory and Performance Optimization of k-d Tree Based Nearest Neighbor Search for 3D Point CloudsReid Pinkham, Shuqing Zeng, Zhengya ZhangHPCA 2020 · 被引用 76 次
相关 Paper
- Fleche: an efficient GPU embedding cache for personalized recommendationsMinhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang 等EuroSys 2022 · 被引用 24 次
- SNARY: A High-Performance and Generic SmartNIC-accelerated Retrieval SystemQiaoyin Gan, Heng Pan, Luyang Li, Kai Lv 等USENIX ATC 2025 · 被引用 2 次
- iMARS: an in-memory-computing architecture for recommendation systemsMengyuan Li, Ann Franchesca Laguna, Dayane Reis, Xunzhao Yin 等DAC 2022 · 被引用 14 次
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 被引用 70 次
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening 等MICRO 2021 · 被引用 31 次
