SKT: A One-Pass Multi-Sketch Data Analytics Accelerator
Monica Chiosa, Thomas Preußer, Gustavo Alonso
Abstract
Data analysts often need to characterize a data stream as a first step to its further processing. Some of the initial insights to be gained include, e.g., the cardinality of the data set and its frequency distribution. Such information is typically extracted by using sketch algorithms, now widely employed to process very large data sets in manageable space and in a single pass over the data. Often, analysts need more than one parameter to characterize the stream. However, computing multiple sketches becomes expensive even when using high-end CPUs. Exploiting the increasing adoption of hardware accelerators, this paper proposes SKT , an FPGA-based accelerator that can compute several sketches along with basic statistics (average, max, min, etc.) in a single pass over the data. SKT has been designed to characterize a data set by calculating its cardinality, its second frequency moment, and its frequency distribution. The design processes data streams coming either from PCIe or TCP/IP, and it is built to fit emerging cloud service architectures, such as Microsoft's Catapult or Amazon's AQUA. The paper explores the trade-offs of designing sketch algorithms on a spatial architecture and how to combine several sketch algorithms into a single design. The empirical evaluation shows how SKT on an FPGA offers a significant performance gain over high-end, server-class CPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3438f9bd-1800-4e82-a79d-a122bb4b052aCited by top-tier papers7
- M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory SystemsYan Sun, Jongyul Kim, Zeduo Yu, Jiyuan Zhang et al.ASPLOS 2025 · 27 citations
- CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation ModelsHailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao et al.SIGMOD 2024 · 15 citations
- Optimistic Data Parallelism for FPGA-Accelerated SketchingMartin Kiefer, Ilias Poulakis, Eleni Tzirita Zacharatou, Volker MarklVLDB 2023 · 11 citations
- TreeSensing: Linearly Compressing Sketches with FlexibilityZirui Liu, Yixin Zhang, Yifan Zhu, Ruwen Zhang et al.SIGMOD 2023 · 10 citations
- CodingSketch: A Hierarchical Sketch with Efficient Encoding and Recursive DecodingQizhi Chen, Yisen Hong, Yuhan Wu, Tong Yang et al.ICDE 2024 · 5 citations
Related papers
- HistSketch: A Compact Data Structure for Accurate Per-Key Distribution MonitoringJintao He, Jiaqi Zhu, Qun HuangICDE 2023 · 21 citations
- OmniSketch: Efficient Multi-Dimensional High-Velocity Stream Analytics with Arbitrary PredicatesWieger R. Punter, Odysseas Papapetrou, Minos N. GarofalakisVLDB 2024 · 10 citations
- Panakos: Chasing the Tails for Multidimensional Data StreamsFuheng Zhao, Punnal Ismail Khan, Divyakant Agrawal, Amr El Abbadi et al.VLDB 2023 · 18 citations
- Spatiotemporal Sketch Disaggregation: Streaming Analytics with Heterogeneous ResourcesJonatan Langlet, Peiqing Chen, Michael Mitzenmacher, Zaoxing Liu et al.ICDE 2026
- Towards Memory-Efficient Streaming Processing with Counter-Cascading Sketching on FPGAMinjin Tang, Mei Wen, Junzhong Shen, Xiaolei Zhao et al.DAC 2020 · 9 citations
