CoopStore: Optimizing Precomputed Summaries for Aggregation
Edward Gan, Peter Bailis, Moses Charikar
Abstract
An emerging class of data systems partition their data and precompute approximate summaries (i.e., sketches and samples) for each segment to reduce query costs. They can then aggregate and combine the segment summaries to estimate results without scanning the raw data. However, given limited storage space each summary introduces approximation errors that affect query accuracy. For instance, systems that use existing mergeable summaries cannot reduce query error below the error of an individual precomputed summary. We introduce Storyboard, a query system that optimizes item frequency and quantile summaries for accuracy when aggregating over multiple segments. Compared to conventional mergeable summaries, Storyboard leverages additional memory available for summary construction and aggregation to derive a more precise combined result. This reduces error by up to 25x over interval aggregations and 4.4x over data cube aggregations on industrial datasets compared to standard summarization methods, with provable worst-case error guarantees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62c309e4-b817-4198-bed6-4d95eb11588fCited by top-tier papers6
- Accelerating Approximate Aggregation Queries with Expensive PredicatesDaniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto et al.VLDB 2021 · 34 citations
- Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query ProcessingXi Liang, Stavros Sintos, Zechao Shang, Sanjay KrishnanSIGMOD 2021 · 27 citations
- Enabling Efficient and General Subpopulation Analytics in Multidimensional Data StreamsAntonis Manousis, Zhuo Cheng, Ran Ben Basat, Zaoxing Liu et al.VLDB 2022 · 16 citations
- Approximate Partition Selection for Big-Data Workloads using Summary StatisticsKexin Rong, Yao Lu, Peter Bailis, Srikanth Kandula et al.VLDB 2020 · 8 citations
- JanusAQP: Efficient Partition Tree Maintenance for Dynamic Approximate Query ProcessingXi Liang, Stavros Sintos, Sanjay KrishnanICDE 2023 · 3 citations
Related papers
- Randomized Sketches for Quantile in LSM-tree based StoreZiling Chen, Shaoxu SongSIGMOD 2025 · 2 citations
- Fast concurrent data sketchesArik Rinberg, Alexander Spiegelman, Edward Bortnikov, Eshcar Hillel et al.PPoPP 2020 · 4 citations
- Sketch-Flip-Merge: Mergeable Sketches for Private Distinct CountingJonathan Hehir, Daniel Ting, Graham CormodeICML 2023 · 12 citations
- FAAQP: Fast and Accurate Approximate Query Processing based on Bitmap-augmented Sum-Product NetworkHanbing Zhang, Yinan Jing, Zhenying He, Kai Zhang et al.SIGMOD 2025
- OmniSketch: Efficient Multi-Dimensional High-Velocity Stream Analytics with Arbitrary PredicatesWieger R. Punter, Odysseas Papapetrou, Minos N. GarofalakisVLDB 2024 · 10 citations
