Cuckoo Heavy Keeper and the balancing act of maintaining heavy hitters in stream processing
Vinh Quang Ngo, Marina Papatriantafilou
Abstract
Finding heavy hitters in databases and data streams is a fundamental problem with applications ranging from network monitoring to database query optimization, machine learning, and more. Approximation algorithms offer practical solutions, but they present tradeoffs involving throughput, memory usage, and accuracy. Moreover, modern applications further complicate these trade-offs by demanding capabilities beyond sequential processing that require both parallel scaling and support for concurrent queries and updates.
Analysis of these trade-offs led us to the key idea behind our proposed streaming algorithm, Cuckoo Heavy Keeper (CHK). The approach introduces an inverted process for distinguishing frequent from infrequent items, which unlocks new algorithmic synergies that were previously inaccessible with conventional approaches. By further analyzing the competing metrics with a focus on parallelism, we propose an algorithmic framework that balances scalability aspects and provides options to optimize query and insertion efficiency based on their relative frequencies. The framework is capable of parallelizing any heavy-hitter detection algorithm.
Besides the algorithms' analysis, we present an extensive evaluation on both real-world and synthetic data across diverse distributions and query selectivity, representing the broad spectrum of application needs. Compared to state-of-the-art methods, CHK improves throughput by 1.7–5.7X and accuracy by up to four orders of magnitude even under low-skew data and tight memory constraints. These properties allow its parallel instances to achieve near-linear scale-up and low latency for heavy-hitter queries, even under a high query rate. We expect the versatility of CHK and its parallel instances to impact a broad spectrum of tools and applications in large-scale data analytics and stream processing systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e68b842e-5b05-4a97-a53f-04c854457ef8Builds on3
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Delegation sketch: a parallel design with support for fast and accurate concurrent operationsCharalampos Stylianopoulos, Ivan Walulya, Magnus Almgren, Olaf Landsiedel et al.EuroSys 2020 · 7 citations
- Fast concurrent data sketchesArik Rinberg, Alexander Spiegelman, Edward Bortnikov, Eshcar Hillel et al.PPoPP 2020 · 4 citations
Related papers
- HeavyLocker: Lock Heavy Hitters in Distributed Data StreamsQilong Shi, Xirui Li, Hanyue Zheng, Tong Yang et al.KDD 2025 · 2 citations
- SketchBuilder: Learning-Augmented Proactive Sketch Construction for Heavy Hitter Detection in Data StreamsYifan Han, Yang Du, Yu-E. Sun, He Huang et al.KDD 2026
- Together is Better: Heavy Hitters Quantile EstimationRana Shahout, Roy Friedman, Ran Ben BasatSIGMOD 2023 · 15 citations
- DUET: A Generic Framework for Finding Special Quadratic Elements in Data StreamsJiaqian Liu, Haipeng Dai, Rui Xia, Meng Li et al.WWW 2022 · 17 citations
- Timely Reporting of Heavy Hitters using External MemoryPrashant Pandey, Shikha Singh, Michael A. Bender, Jonathan W. Berry et al.SIGMOD 2020 · 15 citations
