Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams
Antonis Manousis, Zhuo Cheng, Ran Ben Basat, Zaoxing Liu, Vyas Sekar
摘要
Today's large-scale services ( e.g. , video streaming platforms, data centers, sensor grids) need diverse real-time summary statistics across multiple subpopulations of multidimensional datasets. However, state-of-the-art frameworks do not offer general and accurate analytics in real time at reasonable costs. The root cause is the combinatorial explosion of data subpopulations and the diversity of summary statistics we need to monitor simultaneously. We present Hydra, an efficient framework for multidimensional analytics that presents a novel combination of using a "sketch of sketches" to avoid the overhead of monitoring exponentially-many subpopulations and universal sketching to ensure accurate estimates for multiple statistics. We build Hydra as an Apache Spark plugin and address practical system challenges to minimize overheads at scale. Across multiple real-world and synthetic multidimensional datasets, we show that Hydra can achieve robust error bounds and is an order of magnitude more efficient in terms of operational cost and memory footprint than existing frameworks (e.g., Spark, Druid) while ensuring interactive estimation times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- OmniSketch: Efficient Multi-Dimensional High-Velocity Stream Analytics with Arbitrary PredicatesWieger R. Punter, Odysseas Papapetrou, Minos N. GarofalakisVLDB 2024 · 被引用 10 次
- HeavyLocker: Lock Heavy Hitters in Distributed Data StreamsQilong Shi, Xirui Li, Hanyue Zheng, Tong Yang 等KDD 2025 · 被引用 2 次
- Approximation-First Timeseries Monitoring Query At ScaleZeying Zhu, Jonathan Chamberlain, Kenny Wu, David Starobinski 等VLDB 2025 · 被引用 2 次
- TrustSketch: Trustworthy Sketch-based Telemetry on Cloud HostsZhuo Cheng, Maria Apostolaki, Zaoxing Liu, Vyas SekarNDSS 2024
- CounterSnake: A lossless and generalized compression framework for diverse sketchesXunpeng Liu, Qun Huang, Yaojing Wang, Lihua Miao 等VLDB 2026
它引用的顶会 Paper6
- BeauCoup: Answering Many Network Traffic Queries, One Memory Update at a TimeXiaoqi Chen, Shir Landau Feibish, Mark Braverman, Jennifer RexfordSIGCOMM 2020 · 被引用 91 次
- SALSA: Self-Adjusting Lean Streaming AnalyticsRan Ben Basat, Gil Einziger, Michael Mitzenmacher, Shay VargaftikICDE 2021 · 被引用 45 次
- Faster and More Accurate Measurement through Additive-Error CountersRan Ben Basat, Gil Einziger, Michael Mitzenmacher, Shay VargaftikINFOCOM 2020 · 被引用 17 次
- Joltik: enabling energy-efficient "future-proof" analytics on low-power wide-area networksMingran Yang, Junbo Zhang, Akshay Gadre, Zaoxing Liu 等MobiCom 2020 · 被引用 14 次
- CoopStore: Optimizing Precomputed Summaries for AggregationEdward Gan, Peter Bailis, Moses CharikarVLDB 2020
相关 Paper
- In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them AllRudi Poepsel Lemaitre, Martin Kiefer, Joscha Von Hein, Jorge-Arnulfo Quiané-Ruiz 等VLDB 2021 · 被引用 13 次
- Fast concurrent data sketchesArik Rinberg, Alexander Spiegelman, Edward Bortnikov, Eshcar Hillel 等PPoPP 2020 · 被引用 4 次
- Building Advanced SQL Analytics From Low-Level Plan OperatorsAndré Kohn, Viktor Leis, Thomas NeumannSIGMOD 2021 · 被引用 13 次
- Hyper-USS: Answering Subset Query Over Multi-Attribute Data StreamRuijie Miao, Yiyao Zhang, Guanyu Qu, Kaicheng Yang 等KDD 2023 · 被引用 6 次
- Spatiotemporal Sketch Disaggregation: Streaming Analytics with Heterogeneous ResourcesJonatan Langlet, Peiqing Chen, Michael Mitzenmacher, Zaoxing Liu 等ICDE 2026
