Suda: An Efficient and Secure Unbalanced Data Alignment Framework for Vertical Privacy-Preserving Machine Learning
Lushan Song, Qizhi Zhang, Yu Lin, Haoyu Niu, Daode Zhang, Zheng Qu, Weili Han, Jue Hong, Quanwei Cai, Ye Wu
摘要
Secure data alignment, which securely aligns the data between parties, is the first and crucial step in vertical privacypreserving machine learning (VPPML). Practical applications, e.g. advertising, require VPPML for personalized services. Meanwhile, the data held by parties in these applications are usually unbalanced. Existing secure unbalanced data alignment approaches typically rely on Cuckoo Hashing, which introduces redundant data outside the intersection, leading to significantly increasing communication size during secure training in VPPML. Though secure shuffle operations can trim these redundant data, these operations would incur huge communication overhead. As a result, these secure approaches should be optimized for efficiency in VPPML scenarios. In this paper, we propose Suda, an efficient and secure unbalanced data alignment framework for VPPML. By leveraging polynomial-based operations rather than Cuckoo Hashing, Suda efficiently, directly, and exclusively outputs data shares in the intersection without expensive secure shuffle operations. Consequently, Suda efficiently and seamlessly aligns with secure training in VPPML. Specifically, we first design a novel and efficient batch private information retrieval (PIR) protocol based on the oblivious polynomial reduction and evaluation protocols. Second, we design a batch PIR-to-share protocol extended from the batch PIR protocol with the oblivious polynomial interpolation protocol. Note that the batch PIR-to-share protocol securely obtains feature shares rather than the plaintext features which are the outputs of the batch PIR protocol. Comprehensive experiment results demonstrate that: (1) Suda outperforms the state-of-the-art secure data alignment framework by 31.14× ∼ 210.78× in communication size and up to 8.21× in running time; and (2) Suda outperforms the state-of-the-art batch PIR framework by up to 11.53× in server time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Pisces: Cryptography-based Private Retrieval-Augmented Generation with Dual-Path RetrievalXiaojian Liang, Lushan Song, Shishuai Du, Weicheng Zhu 等ICLR 2026
- Bifrost: A Much Simpler Secure Two-Party Data Join Protocol for Secure Data AnalyticsShuyu Chen, Mingxun Zhou, Haoyu Niu, Guopeng Lin 等VLDB 2026
它引用的顶会 Paper13
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta 等NeurIPS 2021 · 被引用 573 次
- Fast Private Set Intersection from Homomorphic EncryptionHao Chen, Kim Laine, Peter RindalCCS 2017 · 被引用 446 次
- PIR with Compressed Queries and Amortized Query ProcessingSebastian Angel, Hao Chen, Kim Laine, Srinath T. V. SettyS&P 2018 · 被引用 353 次
- Mobile Private Contact Discovery at ScaleDaniel Kales, Christian Rechberger, Thomas Schneider, Matthias Senker 等USENIX Security 2019 · 被引用 157 次
- Blazing Fast PSI from Improved OKVS and Subfield VOLESrinivasan Raghuraman, Peter RindalCCS 2022 · 被引用 81 次
相关 Paper
- Privacy-Preserving Screening for Record LinkageChenyu Huang, Fan Zhang, Huangxun Chen, Yongjun Zhao 等ICDE 2025
- Vectorized Batch Private Information RetrievalMuhammad Haris Mughees, Ling RenS&P 2023
- Batched Differentially Private Information RetrievalKinan Dak Albab, Rawane Issa, Mayank Varia, Kalman GraffiUSENIX Security 2022
- GPU-accelerated PIR with Client-Independent Preprocessing for Large-Scale ApplicationsDaniel Günther, Maurice Heymann, Benny Pinkas, Thomas SchneiderUSENIX Security 2022
- Call Me By My Name: Simple, Practical Private Information Retrieval for Keyword QueriesSofía Celi, Alex DavidsonCCS 2024 · 被引用 13 次
