ParetoES: Hardware-Accelerated Sparse Embedding Similarity via Pareto-Optimal Pruning
Jiaqi Zhai, Xuanhua Shi, Wenju Zhao, Kaiyi Huang, Chencheng Ye, Shunsen Lv, Zhongtian Long, Bingsheng He, Hai Jin
Abstract
Efficient retrieval of sparse embeddings, a critical task in modern information systems, is fundamentally challenged by the memory-bound nature of Top-K sparse matrix-vector multiplication (SpMV). Existing solutions often pursue full computation for absolute accuracy, which is not Pareto-optimal and results in excessive memory transfers and redundant work. We propose ParetoES, an FPGA-accelerated retrieval system that adopts a selective computation strategy to optimize the trade-off between recall and throughput. ParetoES integrates algorithmic and architectural co-design, featuring: (1) a Spherical K-means++ Refine algorithm that combines clustering, low-bit quantization, and unstructured pruning to reduce the candidate search space and memory access overhead; (2) a Hierarchical HotspotBalancing Balance) strategy to mitigate workload skew in multicore environments; and (3) a lightweight Adaptive Cluster Probing Engine (ACPE) architecture with distributed microsorters to enable flexible, high-throughput retrieval. Experiments on five datasets show that when maintaining Recall@100 > 0.8, ParetoES achieves up to and higher Queries Per Second (QPS) than CPU and GPU baselines, respectively. It also demonstrates an average throughput improvement of over the state-of-the-art FPGA accelerator.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b5e19e5e-894c-4f1d-8716-b865460b6e1cRelated papers
- Scaling up HBM Efficiency of Top-K SpMV for Approximate Embedding Similarity on FPGAsAlberto Parravicini, Luca Giuseppe Cellamare, Marco Siracusa, Marco D. SantambrogioDAC 2021 · 18 citations
- AccelES: Accelerating Top-K SpMV for Embedding Similarity via Low-bit PruningJiaqi Zhai, Xuanhua Shi, Kaiyi Huang, Chencheng Ye et al.HPCA 2025 · 2 citations
- SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationHaoqin Huang, Pengcheng Yao, Zhaozeng An, Yufei Sun et al.DAC 2024 · 3 citations
- Scalable top-k retrieval with SpartaGali Sheffi, Dmitry Basin, Edward Bortnikov, David Carmel et al.PPoPP 2020
- Serpens: a high bandwidth memory based accelerator for general-purpose sparse matrix-vector multiplicationLinghao Song, Yuze Chi, Licheng Guo, Jason CongDAC 2022 · 56 citations
