Optimizing Collections of Bloom Filters within a Space Budget
Gabriel Mersy, Zhuo Wang, Stavros Sintos, Sanjay Krishnan
Abstract
With a single Bloom filter, one can approximately answer set membership queries within a space budget. Practical systems often use collections of Bloom filters to facilitate applications such as data skipping, sideways information passing, and network filtering. While the optimal space-to-accuracy allocation is well-understood for a single filter, jointly optimizing how space is used across a collection of filters is yet to be studied. We pose this problem in the following way: (1) let's assume that each Bloom filter has some likelihood of being queried, and (2) given knowledge of this likelihood, how do we allocate space to minimize the expected false positive rate? In other words, "hot" filters are allocated more space, and "cold" filters are allocated less space. In this paper, we show how to solve this optimization problem. We first develop the concept of a "truncated" Bloom filter and theoretically analyze its false positive rate. We then formulate an optimization problem for a collection of truncated Bloom filters that minimizes the false positive rate across a utility distribution while meeting a strict space budget. Next, we show that the problem is convex and find a fast relaxation. Lastly, we apply our method to data skipping and full-text search, demonstrating its effectiveness across the range of possible space budgets when compared to the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f03befd-7a62-4f0c-be17-7f0c0b4fbb4eCited by top-tier papers2
- Mnemosyne: Dynamic Workload-Aware BF Tuning via Accurate Statistics in LSM treesZichen Zhu, Yanpeng Wei, Ju Hyoung Mun, Manos AthanassoulisSIGMOD 2025 · 3 citations
- CounterSnake: A lossless and generalized compression framework for diverse sketchesXunpeng Liu, Qun Huang, Yaojing Wang, Lihua Miao et al.VLDB 2026
Builds on11
- Chucky: A Succinct Cuckoo Filter for LSM-TreeNiv Dayan, Moshe TwittoSIGMOD 2021 · 57 citations
- Stable Learned Bloom Filters for Data StreamsQiyu Liu, Libin Zheng, Yanyan Shen, Lei ChenVLDB 2020 · 45 citations
- Vector Quotient Filters: Overcoming the Time/Space Trade-Off in Filter DesignPrashant Pandey, Alex Conway, Joe Durie, Michael A. Bender et al.SIGMOD 2021 · 40 citations
- InfiniFilter: Expanding Filters to Infinity and BeyondNiv Dayan, Ioana O. Bercea, Pedro Reviriego, Rasmus PaghSIGMOD 2023 · 27 citations
- SplinterDB and Maplets: Improving the Tradeoffs in Key-Value Store Compaction PolicyAlex Conway, Martin Farach-Colton, Rob JohnsonSIGMOD 2023 · 20 citations
Related papers
- Modeling Average False Positive Rates of Recycling Bloom FiltersKahlil Dozier, Loqman Salamatian, Dan RubensteinINFOCOM 2024 · 4 citations
- A four-dimensional Analysis of Partitioned Approximate FiltersTobias Schmidt, Maximilian Bandle, Jana GicevaVLDB 2021 · 6 citations
- Partitioned Learned Bloom FiltersKapil Vaidya, Eric Knorr, Michael Mitzenmacher, Tim KraskaICLR 2021 · 2 citations
- Ensemble Learned Bloom Filters: Two Oracles are Better than OneMing Lin, Lin ChenICML 2025
- Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership QueryingRundong Li, Pinghui Wang, Jiongli Zhu, Junzhou Zhao et al.SIGMOD 2021 · 20 citations
