SelectFormer in Data Markets: Privacy-Preserving and Efficient Data Selection for Transformers with Multi-Party Computation
Xu Ouyang, Felix Xiaozhu Lin, Yangfeng Ji
摘要
Critical to a free data market is private data selection, i.e. the model owner selects and then appraises training data from the data owner before both parties commit to a transaction. To keep the data and model private, this process shall evaluate the target model to be trained over Multi-Party Computation (MPC). While prior work suggests that evaluating Transformer-based models over MPC is prohibitively expensive, this paper makes it practical for the purpose of data selection. Our contributions are three: (1) a new pipeline for private data selection over MPC;
(2) emulating high-dimensional nonlinear operators with low-dimension MLPs, which are trained on a small sample of the data of interest; (3) scheduling MPC in a parallel, multiphase fashion. We evaluate our method on diverse Transformer models and NLP/CV benchmarks. Compared to directly evaluating the target model over MPC, our method reduces the delay from thousands of hours to tens of hours, while only seeing around 0.20% accuracy degradation from training with the selected data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 被引用 2,107 次
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta 等NeurIPS 2021 · 被引用 573 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- QUOTIENT: Two-Party Secure Neural Network Training and PredictionNitin Agrawal, Ali Shahin Shamsabadi, Matt J. Kusner, Adrià GascónCCS 2019 · 被引用 241 次
相关 Paper
- MPC-Pipe: an Efficient Pipeline Scheme for Semi-honest MPC Machine LearningYongqin Wang, Rachit Rajat, Murali AnnavaramASPLOS 2024 · 被引用 5 次
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng 等S&P 2024 · 被引用 149 次
- ABNN2: secure two-party arbitrary-bitwidth quantized neural network predictionsLiyan Shen, Ye Dong, Binxing Fang, Jinqiao Shi 等DAC 2022 · 被引用 11 次
- MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM InferenceWenxuan Zeng, Ye Dong, Jinjin Zhou, Jin Tan 等NeurIPS 2025 · 被引用 4 次
- Privacy-Preserving Feature Selection with Secure Multiparty ComputationXiling Li, Rafael Dowsley, Martine De CockICML 2021 · 被引用 51 次
