Timely Reporting of Heavy Hitters using External Memory
Prashant Pandey, Shikha Singh, Michael A. Bender, Jonathan W. Berry, Martin Farach-Colton, Rob Johnson, Thomas M. Kroeger, Cynthia A. Phillips
Abstract
Given an input stream of size N, a φ-heavy hitter is an item that occurs at least φ N times in S. The problem of finding heavy-hitters is extensively studied in the database literature. We study a real-time heavy-hitters variant in which an element must be reported shortly after we see its T = φ N-th occurrence (and hence becomes a heavy hitter). We call this the Timely Event Detection (TED) Problem. The TED problem models the needs of many real-world monitoring systems, which demand accurate (i.e., no false negatives) and timely reporting of all events from large, high-speed streams, and with a low reporting threshold (high sensitivity). Like the classic heavy-hitters problem, solving the TED problem without false-positives requires large space (Ω(N) words). Thus in-RAM heavy-hitters algorithms typically sacrifice accuracy (i.e., allow false positives), sensitivity, or timeliness (i.e., use multiple passes). We show how to adapt heavy-hitters algorithms to external memory to solve the TED problem on large high-speed streams while guaranteeing accuracy, sensitivity, and timeliness. Our data structures are limited only by I/O-bandwidth (not latency) and support a tunable trade-off between reporting delay and I/O overhead. With a small bounded reporting delay, our algorithms incur only a logarithmic I/O overhead. We implement and validate our data structures empirically using the Firehose streaming benchmark. Multi-threaded versions of our structures can scale to process 11M observations per second before becoming CPU bound. In comparison, a naive adaptation of the standard heavy-hitters algorithm to external memory would be limited by the storage device's random I/O throughput, i.e., 100K observations per second.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d8666bf-c792-4315-9082-7a594fcd3a4bCited by top-tier papers4
- Vector Quotient Filters: Overcoming the Time/Space Trade-Off in Filter DesignPrashant Pandey, Alex Conway, Joe Durie, Michael A. Bender et al.SIGMOD 2021 · 40 citations
- Hyper-USS: Answering Subset Query Over Multi-Attribute Data StreamRuijie Miao, Yiyao Zhang, Guanyu Qu, Kaicheng Yang et al.KDD 2023 · 6 citations
- Zombie Hashing: Reanimating Tombstones in GraveyardYuvaraj Chesetti, Benwei Shi, Jeff M. Phillips, Prashant PandeySIGMOD 2025 · 2 citations
- HeavyLocker: Lock Heavy Hitters in Distributed Data StreamsQilong Shi, Xirui Li, Hanyue Zheng, Tong Yang et al.KDD 2025 · 2 citations
Related papers
- Cuckoo Heavy Keeper and the balancing act of maintaining heavy hitters in stream processingVinh Quang Ngo, Marina PapatriantafilouVLDB 2025 · 2 citations
- SketchBuilder: Learning-Augmented Proactive Sketch Construction for Heavy Hitter Detection in Data StreamsYifan Han, Yang Du, Yu-E. Sun, He Huang et al.KDD 2026
- Together is Better: Heavy Hitters Quantile EstimationRana Shahout, Roy Friedman, Ran Ben BasatSIGMOD 2023 · 15 citations
- Persistent Items Tracking in Large Data Streams Based on Adaptive SamplingLin Chen, Raphael C.-W. Phan, Zhili Chen, Dan HuangINFOCOM 2022 · 15 citations
- Information Based Heavy Hitters for Real-Time DNS Data Exfiltration DetectionYarin Ozery, Asaf Nadler, Asaf ShabtaiNDSS 2024
