Clipped SGD Algorithms for Performative Prediction: Tight Bounds for Stochastic Bias and Remedies
Qiang Li, Michal Yemini, Hoi-To Wai
Abstract
This paper studies the convergence of clipped stochastic gradient descent (SGD) algorithms with decision-dependent data distribution. Our setting is motivated by privacy preserving optimization algorithms that interact with performative data where the prediction models can influence future outcomes. This challenging setting involves the non-smooth clipping operator and non-gradient dynamics due to distribution shifts. We make two contributions in pursuit for a performative stable solution using clipped SGD algorithms. First, we characterize the clipping bias with projected clipped SGD (PCSGD) algorithm which is caused by the clipping operator that prevents PCSGD from reaching a stable solution. When the loss function is strongly convex, we quantify the lower and upper bounds for this clipping bias and demonstrate a bias amplification phenomenon with the sensitivity of data distribution. When the loss function is non-convex, we bound the magnitude of stationarity bias. Second, we propose remedies to mitigate the bias either by utilizing an optimal step size design for PCSGD, or to apply the recent DiceSGD algorithm [Zhang et al., 2024] . Our analysis is also extended to show that the latter algorithm is free from clipping bias in the performative setting. Numerical experiments verify our findings. * M. Yemini is with the faculty of Engineering, Bar-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cf32888-74d2-4ecb-897a-2e5f5f0e3147Builds on14
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Performative PredictionJuan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz HardtICML 2020 · 422 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- Understanding Gradient Clipping in Private SGD: A Geometric PerspectiveXiangyi Chen, Zhiwei Steven Wu, Mingyi HongNeurIPS 2020 · 254 citations
Related papers
- Stochastic Optimization Schemes for Performative Prediction with Nonconvex LossQiang Li, Hoi-To WaiNeurIPS 2024 · 18 citations
- Improved Convergence of Differential Private SGD with Gradient ClippingHuang Fang, Xiaoyun Li, Chenglin Fan, Ping LiICLR 2023
- Eliminating Solution Bias in Differentially Private OptimizationDONGRUN LI, YUN ZENG, Zibo Wei, Jiacheng Wei et al.ICML 2026
- Differentially Private SGD Without Clipping Bias: An Error-Feedback ApproachXinwei Zhang, Zhiqi Bu, Steven Wu, Mingyi HongICLR 2024 · 15 citations
- A Theory to Instruct Differentially-Private Learning via Clipping Bias ReductionHanshen Xiao, Zihang Xiang, Di Wang, Srinivas DevadasS&P 2023
