Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams
Antonis Manousis, Zhuo Cheng, Ran Ben Basat, Zaoxing Liu, Vyas Sekar
Abstract
Today's large-scale services ( e.g. , video streaming platforms, data centers, sensor grids) need diverse real-time summary statistics across multiple subpopulations of multidimensional datasets. However, state-of-the-art frameworks do not offer general and accurate analytics in real time at reasonable costs. The root cause is the combinatorial explosion of data subpopulations and the diversity of summary statistics we need to monitor simultaneously. We present Hydra, an efficient framework for multidimensional analytics that presents a novel combination of using a "sketch of sketches" to avoid the overhead of monitoring exponentially-many subpopulations and universal sketching to ensure accurate estimates for multiple statistics. We build Hydra as an Apache Spark plugin and address practical system challenges to minimize overheads at scale. Across multiple real-world and synthetic multidimensional datasets, we show that Hydra can achieve robust error bounds and is an order of magnitude more efficient in terms of operational cost and memory footprint than existing frameworks (e.g., Spark, Druid) while ensuring interactive estimation times.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bce6cc98-a6ed-44f2-8ff6-27aa3704d967Cited by top-tier papers6
- OmniSketch: Efficient Multi-Dimensional High-Velocity Stream Analytics with Arbitrary PredicatesWieger R. Punter, Odysseas Papapetrou, Minos N. GarofalakisVLDB 2024 · 10 citations
- HeavyLocker: Lock Heavy Hitters in Distributed Data StreamsQilong Shi, Xirui Li, Hanyue Zheng, Tong Yang et al.KDD 2025 · 2 citations
- Approximation-First Timeseries Monitoring Query At ScaleZeying Zhu, Jonathan Chamberlain, Kenny Wu, David Starobinski et al.VLDB 2025 · 2 citations
- TrustSketch: Trustworthy Sketch-based Telemetry on Cloud HostsZhuo Cheng, Maria Apostolaki, Zaoxing Liu, Vyas SekarNDSS 2024
- CounterSnake: A lossless and generalized compression framework for diverse sketchesXunpeng Liu, Qun Huang, Yaojing Wang, Lihua Miao et al.VLDB 2026
Builds on6
- BeauCoup: Answering Many Network Traffic Queries, One Memory Update at a TimeXiaoqi Chen, Shir Landau Feibish, Mark Braverman, Jennifer RexfordSIGCOMM 2020 · 91 citations
- SALSA: Self-Adjusting Lean Streaming AnalyticsRan Ben Basat, Gil Einziger, Michael Mitzenmacher, Shay VargaftikICDE 2021 · 45 citations
- Faster and More Accurate Measurement through Additive-Error CountersRan Ben Basat, Gil Einziger, Michael Mitzenmacher, Shay VargaftikINFOCOM 2020 · 17 citations
- Joltik: enabling energy-efficient "future-proof" analytics on low-power wide-area networksMingran Yang, Junbo Zhang, Akshay Gadre, Zaoxing Liu et al.MobiCom 2020 · 14 citations
- CoopStore: Optimizing Precomputed Summaries for AggregationEdward Gan, Peter Bailis, Moses CharikarVLDB 2020
Related papers
- In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them AllRudi Poepsel Lemaitre, Martin Kiefer, Joscha Von Hein, Jorge-Arnulfo Quiané-Ruiz et al.VLDB 2021 · 13 citations
- Fast concurrent data sketchesArik Rinberg, Alexander Spiegelman, Edward Bortnikov, Eshcar Hillel et al.PPoPP 2020 · 4 citations
- Building Advanced SQL Analytics From Low-Level Plan OperatorsAndré Kohn, Viktor Leis, Thomas NeumannSIGMOD 2021 · 13 citations
- Hyper-USS: Answering Subset Query Over Multi-Attribute Data StreamRuijie Miao, Yiyao Zhang, Guanyu Qu, Kaicheng Yang et al.KDD 2023 · 6 citations
- Spatiotemporal Sketch Disaggregation: Streaming Analytics with Heterogeneous ResourcesJonatan Langlet, Peiqing Chen, Michael Mitzenmacher, Zaoxing Liu et al.ICDE 2026
