Complaint-Driven Training Data Debugging at Interactive Speeds
Lampros Flokas, Weiyuan Wu, Yejia Liu, Jiannan Wang, Nakul Verma, Eugene Wu
摘要
Modern databases support queries that perform model inference (inference queries). Although powerful and widely used, inference queries are susceptible to incorrect results if the model is biased due to training data errors. Recently, prior work Rain proposed complaint-driven data debugging which uses user-specified errors in the output of inference queries (Complaints) to rank erroneous training examples that most likely caused the complaint. This can help users better interpret results and debug training sets. Rain combined influence analysis from the ML literature with relaxed query provenance polynomials from the DB literature to approximate the derivative of complaints w.r.t. training examples. Although effective, the runtime is O(|T|d), where T and d are the training set and model sizes, due to its reliance on the model's second order derivatives (the Hessian). On a Wide Resnet Network (WRN) model with 1.5 million parameters, it takes >1 minute to debug a complaint. We observe that most complaint debugging costs are independent of the complaint, and that modern models are overparameterized. In response, Rain++ uses precomputation techniques, based on non-trivial insights unique to data debugging, to reduce debugging latencies to a constant factor independent of model size. We also develop optimizations when the queried database is known apriori, and for standing queries over streaming databases. Combining these optimizations in Rain++ ensures interactive debugging latencies ( 1ms) on models with millions of parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami 等NeurIPS 2020 · 被引用 434 次
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 被引用 262 次
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 被引用 36 次
- PrIU: A Provenance-Based Approach for Incrementally Updating Regression ModelsYinjun Wu, Val Tannen, Susan B. DavidsonSIGMOD 2020 · 被引用 26 次
相关 Paper
- Enabling SQL-based Training Data Debugging for Federated LearningYejia Liu, Weiyuan Wu, Lampros Flokas, Jiannan Wang 等VLDB 2022 · 被引用 16 次
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Facilitating SQL Query Composition and AnalysisZainab Zolaktaf, Mostafa Milani, Rachel PottingerSIGMOD 2020 · 被引用 19 次
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 被引用 54 次
- FaDE: More Than a Million What-ifs Per SecondHaneen Mohammed, Eugene Wu, Alexander Yao, Charlie Summers 等VLDB 2025 · 被引用 6 次
