Differential Privacy Under Class Imbalance: Methods and Empirical Insights
Lucas Rosenblatt, Yuliia Lut, Ethan Turok, Marco Avella Medina, Rachel Cummings
摘要
Imbalanced learning occurs in classification settings where the distribution of class-labels is highly skewed in the training data, such as when predicting rare diseases or in fraud detection. This class imbalance presents a significant algorithmic challenge, which can be further exacerbated when privacypreserving techniques such as differential privacy are applied to protect sensitive training data. Our work formalizes these challenges and provides a number of algorithmic solutions. We consider DP variants of pre-processing methods that privately augment the original dataset to reduce the class imbalance; these include oversampling, SMOTE, and private synthetic data generation. We also consider DP variants of in-processing techniques, which adjust the learning algorithm to account for the imbalance; these include model bagging, class-weighted empirical risk minimization and class-weighted deep learning. For each method, we either adapt an existing imbalanced learning technique to the private setting or demonstrate its incompatibility with differential privacy. Finally, we empirically evaluate these privacy-preserving imbalanced learning methods under various data and distributional settings. We find that private synthetic data methods perform well as a data pre-processing step, while class-weighted ERMs are an alternative in higher-dimensional settings where private synthetic data suffers from the curse of dimensionality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority OversamplingGeorgi Ganev, MohammadReza Nazari, Rees Davison, Amirhassan Fallah Dizche 等ICLR 2026 · 被引用 6 次
- Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular DataLucas Rosenblatt, Peihan Liu, Ryan McKenna, Natalia PonomarevaICML 2026 · 被引用 1 次
它引用的顶会 Paper18
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Rethinking the Value of Labels for Improving Class-Imbalanced LearningYuzhe Yang, Zhi XuNeurIPS 2020 · 被引用 512 次
- Hyperparameter Tuning with Renyi Differential PrivacyNicolas Papernot, Thomas SteinkeICLR 2022 · 被引用 157 次
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
相关 Paper
- Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic DataGeorgi Ganev, Bristena Oprisanu, Emiliano De CristofaroICML 2022 · 被引用 78 次
- Differentially Private Prototypes for Imbalanced Transfer LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischAAAI 2025 · 被引用 4 次
- INO-SGD: Addressing Utility Imbalance under Individualized Differential PrivacyXiao Tian, Jue Fan, Rachael Hwee Ling Sim, Bryan Kian Hsiang LowICLR 2026
- Synthetic Tabular Data Generation for Imbalanced Classification: The Surprising Effectiveness of an Overlap ClassAnnie D'souza, Swetha M, Sunita SarawagiAAAI 2025 · 被引用 9 次
- Multi-Class Support Vector Machine with Differential PrivacyJinseong Park, Yujin Choi, Jaewook LeeNeurIPS 2025 · 被引用 1 次
