FaDE: More Than a Million What-ifs Per Second
Haneen Mohammed, Eugene Wu, Alexander Yao, Charlie Summers, Lampros Flokas, Gromit Yeuk-Yin Chan, Subrata Mitra, Hongbin Zhong
Abstract
What-if queries are the building blocks for many explanation and analytics applications—sensitivity analysis, hypothetical reasoning, data cleaning, probabilistic databases—that explore how a query's output changes due to input data changes. Their response time is bounded by intervention evaluation latency, which can be in the minute or hours for complex queries and large datasets. FaDE is a compilation engine that uses provenance to evaluate hypothetical deletion and scaling interventions at low latency and high throughput. FaDE forgoes conventional provenance representations as symbolic expressions and leverages their underlying relational structure. This accelerates intervention evaluation on average by 1000× against IVM and 10,000× against prior provenance-based approaches. In addition, FaDE develops a suite of optimizations (e.g., compilation, parallelization, incremental evaluation, sparse representations) that collectively raise evaluation throughput to >1 million interventions per sec—a rate that can brute-force existing applications within 1 s.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f842f11b-2039-4aa5-a397-5eebcbf3ec92Builds on3
- Computing How-Provenance for SPARQL Queries via Query RewritingDaniel Hernández, Luis Galárraga, Katja HoseVLDB 2021 · 42 citations
- Provenance-based Data SkippingXing Niu, Boris Glavic, Ziyu Liu, Pengyuan Li et al.VLDB 2022 · 9 citations
- ProvTalk: Towards Interpretable Multi-level Provenance Analysis in Networking Functions Virtualization (NFV)Azadeh Tabiban, Heyang Zhao, Yosr Jarraya, Makan Pourzandi et al.NDSS 2022
Related papers
- Complaint-Driven Training Data Debugging at Interactive SpeedsLampros Flokas, Weiyuan Wu, Yejia Liu, Jiannan Wang et al.SIGMOD 2022 · 13 citations
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis]Yeounoh Chung, Rushabh Desai, Jian He, Yu Xiao et al.SIGMOD 2026 · 8 citations
- Computing the Shapley Value of Facts in Query AnsweringDaniel Deutch, Nave Frost, Benny Kimelfeld, Mikaël MonetSIGMOD 2022 · 31 citations
- CloudGlide: Deconstructing the Landscape of Cloud-Based AnalyticsMichail Georgoulakis Misegiannis, Daniel Ritter, Viktor Leis, Jana GicevaVLDB 2025 · 1 citation
- Toward Temporal Attribution Analytics in Dataflows [Vision Paper]Chrysanthi Kosyfaki, Ruiyuan Zhang, Nikos Mamoulis, Xiaofang ZhouVLDB 2026
