Enabling SQL-based Training Data Debugging for Federated Learning
Yejia Liu, Weiyuan Wu, Lampros Flokas, Jiannan Wang, Eugene Wu
摘要
How can we debug a logistic regression model in a federated learning setting when seeing the model behave unexpectedly (e.g., the model rejects all high-income customers' loan applications)? The SQL-based training data debugging framework has proved effective to fix this kind of issue in a non-federated learning setting. Given an unexpected query result over model predictions, this framework automatically removes the label errors from training data such that the unexpected behavior disappears in the retrained model. In this paper, we enable this powerful framework for federated learning. The key challenge is how to develop a security protocol for federated debugging which is proved to be secure, efficient, and accurate. Achieving this goal requires us to investigate how to seamlessly integrate the techniques from multiple fields (Databases, Machine Learning, and Cybersecurity). We first propose FedRain, which extends Rain, the state-of-the-art SQL-based training data debugging framework, to our federated learning setting. We address several technical challenges to make FedRain work and analyze its security guarantee and time complexity. The analysis results show that FedRain falls short in terms of both efficiency and security. To overcome these limitations, we redesign our security protocol and propose Frog, a novel SQL-based training data debugging framework tailored for federated learning. Our theoretical analysis shows that Frog is more secure, more accurate, and more efficient than FedRain. We conduct extensive experiments using several real-world datasets and a case study. The experimental results are consistent with our theoretical analysis and validate the effectiveness of Frog in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Falcon: A Privacy-Preserving and Interpretable Vertical Federated Learning SystemYuncheng Wu, Naili Xing, Gang Chen, Tien Tuan Anh Dinh 等VLDB 2023 · 被引用 47 次
- FEAST: A Communication-efficient Federated Feature Selection Framework for Relational DataRui Fu, Yuncheng Wu, Quanqing Xu, Meihui ZhangSIGMOD 2023 · 被引用 17 次
- Secure and Verifiable Data Collaboration with Low-Cost Zero-Knowledge ProofsYizheng Zhu, Yuncheng Wu, Zhaojing Luo, Beng Chin Ooi 等VLDB 2024 · 被引用 14 次
它引用的顶会 Paper6
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Fast Private Set Intersection from Homomorphic EncryptionHao Chen, Kim Laine, Peter RindalCCS 2017 · 被引用 446 次
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen 等VLDB 2020 · 被引用 259 次
- Practical Federated Gradient Boosting Decision TreesQinbin Li, Zeyi Wen, Bingsheng HeAAAI 2020 · 被引用 215 次
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 被引用 36 次
相关 Paper
- FedDebug: Systematic Debugging for Federated Learning ApplicationsWaris Gill, Ali Anwar, Muhammad Ali GulzarICSE 2023 · 被引用 13 次
- Efficient Federated-Learning Model DebuggingAnran Li, Lan Zhang, Junhao Wang, Juntao Tan 等ICDE 2021 · 被引用 37 次
- Complaint-Driven Training Data Debugging at Interactive SpeedsLampros Flokas, Weiyuan Wu, Yejia Liu, Jiannan Wang 等SIGMOD 2022 · 被引用 13 次
- Generative Models for Effective ML on Private, Decentralized DatasetsSean Augenstein, H. Brendan McMahan, Daniel Ramage, Swaroop Ramaswamy 等ICLR 2020 · 被引用 207 次
- Understanding the Bug Characteristics and Fix Strategies of Federated Learning SystemsXiaohu Du, Xiao Chen, Jialun Cao, Ming Wen 等FSE 2023 · 被引用 6 次
