Lessons from the AdKDD'21 Privacy-Preserving ML Challenge
Eustache Diemert, Romain Fabre, Alexandre Gilotte, Fei Jia, Basile Leparmentier, Jérémie Mary, Zhonghua Qu, Ugo Tanielian, Hui Yang
Abstract
Designing data sharing mechanisms providing performance and strong privacy guarantees is a hot topic for the Online Advertising industry. Namely, a prominent proposal discussed under the Improving Web Advertising Business Group at W3C only allows sharing advertising signals through aggregated, differentially private reports of past displays. To study this proposal extensively, an open Privacy-Preserving Machine Learning Challenge took place at AdKDD'21, a premier workshop on Advertising Science with data provided by advertising company Criteo. In this paper, we describe the challenge tasks, the structure of the available datasets, report the challenge results, and enable its full reproducibility. A key finding is that learning models on large, aggregated data in the presence of a small set of unaggregated data points can be surprisingly efficient and cheap. We also run additional experiments to observe the sensitivity of winning methods to different parameters such as privacy budget or quantity of available privileged side information. We conclude that the industry needs either alternate designs for private data sharing or a breakthrough in learning with aggregated data only to keep ad relevance at a reasonable level. CCS CONCEPTS • Information systems → Online advertising; • Security and privacy → Economics of security and privacy; • Computing methodologies → Learning settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dfef411-b354-4529-b287-e0f324db5c88Cited by top-tier papers3
- Label differential privacy and private training data releaseRóbert Istvan Busa-Fekete, Andrés Muñoz Medina, Umar Syed, Sergei VassilvitskiiICML 2023 · 9 citations
- Learning from Label Proportions via Proportional Value ClassificationTianhao Ma, Wei Wang, Ximing Li, Gang Niu et al.ICLR 2026
- Particle Flow for Learning from Label Proportionsalain rakotomamonjy, Maxime Vono, Ralaivola LivaICML 2026
Builds on2
Related papers
- Learning from Aggregate responses: Instance Level versus Bag Level Loss FunctionsAdel Javanmard, Lin Chen, Vahab Mirrokni, Ashwinkumar Badanidiyuru et al.ICLR 2024 · 3 citations
- The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM modelAditya Desai, Anshumali ShrivastavaNeurIPS 2022 · 17 citations
- Privacy Budgeting for Growing Machine Learning DatasetsWeiting Li, Liyao Xiang, Zhou Zhou, Feng PengINFOCOM 2021 · 14 citations
- A Protocol for Privately Reporting Ad Impressions at ScaleMatthew Green, Watson Ladd, Ian MiersCCS 2016 · 78 citations
- The Value of Collaboration in Convex Machine Learning with Differential PrivacyNan Wu, Farhad Farokhi, David B. Smith, Mohamed Ali KâafarS&P 2020 · 112 citations
