Scaling up HBM Efficiency of Top-K SpMV for Approximate Embedding Similarity on FPGAs
Alberto Parravicini, Luca Giuseppe Cellamare, Marco Siracusa, Marco D. Santambrogio
摘要
Top-K SpMV is a key component of similarity-search on sparse embeddings. This sparse workload does not perform well on general-purpose NUMA systems that employ traditional caching strategies. Instead, modern FPGA accelerator cards have a few tricks up their sleeve. We introduce a Top-KSpMV FPGA design that leverages reduced precision and a novel packet-wise CSR matrix compression, enabling custom data layouts and delivering bandwidth efficiency often unreachable even in architectures with higher peak bandwidth. With HBM-based boards, we are 100x faster than a multi-threaded CPU implementation and 2x faster than a GPU with 20% higher bandwidth, with 14.2x higher power-efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- ParetoES: Hardware-Accelerated Sparse Embedding Similarity via Pareto-Optimal PruningJiaqi Zhai, Xuanhua Shi, Wenju Zhao, Kaiyi Huang 等ISCA 2026
- AccelES: Accelerating Top-K SpMV for Embedding Similarity via Low-bit PruningJiaqi Zhai, Xuanhua Shi, Kaiyi Huang, Chencheng Ye 等HPCA 2025 · 被引用 2 次
- Serpens: a high bandwidth memory based accelerator for general-purpose sparse matrix-vector multiplicationLinghao Song, Yuze Chi, Licheng Guo, Jason CongDAC 2022 · 被引用 56 次
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari 等HPCA 2021 · 被引用 66 次
