Cocoon: A System Architecture for Differentially Private Training with Correlated Noises
Donghwan Kim, Xin Gu, Jinho Baek, Timothy Lo, Younghoon Min, Kwangsik Shin, Jongryool Kim, Jongse Park, Kiwan Maeng
Abstract
Machine learning (ML) models memorize and leak training data, causing serious privacy issues to data owners. Training algorithms with differential privacy (DP) have been gaining attention as a solution. However, these algorithms add noise at each training iteration and degrade accuracy, limiting their real-world adoption. To improve accuracy, a new family of approaches adds carefully designed correlated noises, so that noises cancel out each other across iterations. We performed an extensive characterization study of these new mechanisms and show they incur non-negligible overheads when the model is relatively large or uses large embedding tables compared to the hardware capacity. Motivated by the analysis, we propose Cocoon, a framework for efficient training with correlated noises. Cocoon stores and processes the large noise history across CPU, GPU, and memory extension module, introduces optimizations for sparse embedding tables, and leverages to-be-commercialized near-memory processing (NMP) devices. On a real system with an FPGA-based NMP device prototype, Cocoon improves the performance by 1.23–10.82×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60158c7f-0bbd-43e7-96e9-b679d6957ba6Builds on53
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated LearningMilad Nasr, Reza Shokri, Amir HoumansadrS&P 2019 · 1,778 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Differentially Private Model Publishing for Deep LearningLei Yu, Ling Liu, Calton Pu, Mehmet Emre Gursoy et al.S&P 2019 · 294 citations
- Differentially Private Prototypes for Imbalanced Transfer LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischAAAI 2025 · 4 citations
- DiVa: An Accelerator for Differentially Private Machine LearningBeomsik Park, Ranggi Hwang, Dongho Yoon, Yoonhyuk Choi et al.MICRO 2022 · 12 citations
- All Rivers Run to the Sea: Private Learning with Asymmetric FlowsYue Niu, Ramy E. Ali, Saurav Prakash, Salman AvestimehrCVPR 2024
- DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise ReductionXinwei Zhang, Zhiqi Bu, Borja Balle, Mingyi Hong et al.ICLR 2025
