PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models
Yinjun Wu, Val Tannen, Susan B. Davidson
Abstract
The ubiquitous use of machine learning algorithms brings new challenges to traditional database problems such as incremental view update. Much effort is being put in better understanding and debugging machine learning models, as well as in identifying and repairing errors in training datasets. Our focus is on how to assist these activities when they have to retrain the machine learning model after removing problematic training samples in cleaning or selecting different subsets of training data for interpretability. This paper presents an efficient provenance-based approach, PrIU, and its optimized version, PrIU-opt, for incrementally updating model parameters without sacrificing prediction accuracy. We prove the correctness and convergence of the incrementally updated model parameters, and validate it experimentally. Experimental results show that up to two orders of magnitude speed-ups can be achieved by PrIU-opt compared to simply retraining the model from scratch, yet obtaining highly similar models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a37d1bde-2d0a-4a7a-88c1-5314e48f92f6Cited by top-tier papers9
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 262 citations
- Federated Unlearning via Class-Discriminative PruningJunxiao Wang, Song Guo, Xin Xie, Heng QiWWW 2022 · 217 citations
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 53 citations
- Capturing and querying fine-grained provenance of preprocessing pipelines in data scienceAdriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo TorloneVLDB 2021 · 39 citations
- Complaint-Driven Training Data Debugging at Interactive SpeedsLampros Flokas, Weiyuan Wu, Yejia Liu, Jiannan Wang et al.SIGMOD 2022 · 13 citations
Builds on1
Related papers
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 36 citations
- Detect, Distill and Update: Learned DB Systems Facing Out of Distribution DataMeghdad Kurmanji, Peter TriantafillouSIGMOD 2023 · 19 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 13 citations
- DistVec: Efficient Distributed Machine Learning in Parallel Database SystemsXinyi Zhang, Liangzu Liu, Xupeng Miao, Yinjun Wu et al.ICDE 2026
