Lune

DAC2021Top-tier venue

Scaling up HBM Efficiency of Top-K SpMV for Approximate Embedding Similarity on FPGAs

Alberto Parravicini, Luca Giuseppe Cellamare, Marco Siracusa, Marco D. Santambrogio

2021Year
18Citations

Abstract

Top-K SpMV is a key component of similarity-search on sparse embeddings. This sparse workload does not perform well on general-purpose NUMA systems that employ traditional caching strategies. Instead, modern FPGA accelerator cards have a few tricks up their sleeve. We introduce a Top-KSpMV FPGA design that leverages reduced precision and a novel packet-wise CSR matrix compression, enabling custom data layouts and delivering bandwidth efficiency often unreachable even in architectures with higher peak bandwidth. With HBM-based boards, we are 100x faster than a multi-threaded CPU implementation and 2x faster than a GPU with 20% higher bandwidth, with 14.2x higher power-efficiency.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b55ed807-fb47-4258-ad36-e52a4e02f145

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines