Complaint-driven Training Data Debugging for Query 2.0
Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan Wang
Abstract
As the need for machine learning (ML) increases rapidly across all industry sectors, there is a significant interest among commercial database providers to support "Query 2.0", which integrates model inference into SQL queries. Debugging Query 2.0 is very challenging since an unexpected query result may be caused by the bugs in training data (e.g., wrong labels, corrupted features). In response, we propose Rain, a complaint-driven training data debugging system. Rain allows users to specify complaints over the query's intermediate or final output, and aims to return a minimum set of training examples so that if they were removed, the complaints would be resolved. To the best of our knowledge, we are the first to study this problem. A naive solution requires retraining an exponential number of ML models. We propose two novel heuristic approaches based on influence functions which both require linear retraining steps. We provide an in-depth analytical and empirical analysis of the two approaches and conduct extensive experiments to evaluate their effectiveness using four real-world datasets. Results show that Rain achieves the highest [email protected] among all the baselines while still returns results interactively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Measuring the Effect of Training Data on Deep Learning Predictions via Randomized ExperimentsJinkun Lin, Anqi Zhang, Mathias Lécuyer, Jinyang Li et al.ICML 2022 · 70 citations
- MAAT: a novel ensemble approach to addressing fairness and performance bugs for machine learning softwareZhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2022 · 65 citations
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang et al.VLDB 2021 · 55 citations
- Hyper-Tune: Towards Efficient Hyper-parameter Tuning at ScaleYang Li, Yu Shen, Huaijun Jiang, Wentao Zhang et al.VLDB 2022 · 32 citations
- XInsight: eXplainable Data Analysis Through The Lens of CausalityPingchuan Ma, Rui Ding, Shuai Wang, Shi Han et al.SIGMOD 2023 · 20 citations
Related papers
- Complaint-Driven Training Data Debugging at Interactive SpeedsLampros Flokas, Weiyuan Wu, Yejia Liu, Jiannan Wang et al.SIGMOD 2022 · 13 citations
- Enabling SQL-based Training Data Debugging for Federated LearningYejia Liu, Weiyuan Wu, Lampros Flokas, Jiannan Wang et al.VLDB 2022 · 16 citations
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- PrIU: A Provenance-Based Approach for Incrementally Updating Regression ModelsYinjun Wu, Val Tannen, Susan B. DavidsonSIGMOD 2020 · 26 citations
