Lune

ISCA2026Top-tier venue

ParetoES: Hardware-Accelerated Sparse Embedding Similarity via Pareto-Optimal Pruning

Jiaqi Zhai, Xuanhua Shi, Wenju Zhao, Kaiyi Huang, Chencheng Ye, Shunsen Lv, Zhongtian Long, Bingsheng He, Hai Jin

2026Year

Abstract

Efficient retrieval of sparse embeddings, a critical task in modern information systems, is fundamentally challenged by the memory-bound nature of Top-K sparse matrix-vector multiplication (SpMV). Existing solutions often pursue full computation for absolute accuracy, which is not Pareto-optimal and results in excessive memory transfers and redundant work. We propose ParetoES, an FPGA-accelerated retrieval system that adopts a selective computation strategy to optimize the trade-off between recall and throughput. ParetoES integrates algorithmic and architectural co-design, featuring: (1) a Spherical K-means++ Refine algorithm that combines clustering, low-bit quantization, and unstructured pruning to reduce the candidate search space and memory access overhead; (2) a Hierarchical HotspotBalancing (H2(H^{2} Balance) strategy to mitigate workload skew in multicore environments; and (3) a lightweight Adaptive Cluster Probing Engine (ACPE) architecture with distributed microsorters to enable flexible, high-throughput retrieval. Experiments on five datasets show that when maintaining Recall@100 > 0.8, ParetoES achieves up to 540×\mathbf{5 4 0} \times and 79×\mathbf{7 9} \times higher Queries Per Second (QPS) than CPU and GPU baselines, respectively. It also demonstrates an average throughput improvement of 2.27×2.27 \times over the state-of-the-art FPGA accelerator.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get b5e19e5e-894c-4f1d-8716-b865460b6e1c

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines