GPU-based Private Information Retrieval for On-Device Machine Learning Inference
Maximilian Lam, Jeff Johnson, Wenjie Xiong, Kiwan Maeng, Udit Gupta, Yang Li, Liangzhen Lai, Ilias Leontiadis, Minsoo Rhu, Hsien-Hsin S. Lee, Vijay Janapa Reddi, Gu-Yeon Wei
摘要
On-device machine learning (ML) inference can enable the use of private user data on user devices without revealing them to remote servers. However, a pure on-device solution to private ML inference is impractical for many applications that rely on embedding tables that are too large to be stored on-device. In particular, recommendation models typically use multiple embedding tables each on the order of 1--10 GBs of data, making them impractical to store on-device. To overcome this barrier, we propose the use of private information retrieval (PIR) to efficiently and privately retrieve embeddings from servers without sharing any private information. As off-the-shelf PIR algorithms are usually too computationally intensive to directly use for latency-sensitive inference tasks, we 1) propose novel GPU-based acceleration of PIR, and 2) co-design PIR with the downstream ML application to obtain further speedup. Our GPU acceleration strategy improves system throughput by more than 20× over an optimized CPU PIR implementation, and our PIR-ML co-design provides an over 5× additional throughput improvement at fixed model quality. Together, for various on-device ML applications such as recommendation and language modeling, our system on a single V100 GPU can serve up to 100,000 queries per second---a > 100× throughput improvement over a CPU-based baseline---while maintaining model accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory ProcessingChenqi Lin, Kang Yang, Tianshi Xu, Ling Liang 等MICRO 2025 · 被引用 4 次
- Practical Federated Recommendation Model Learning Using ORAM with Controlled PrivacyJinyu Liu, Wenjie Xiong, G. Edward Suh, Kiwan MaengASPLOS 2025 · 被引用 2 次
- Do You Need a Receipt? Anonymous Credential Revocation at Continental Scale via Private Record CertificationKasra EdalatNejad, Sebastian Faust, Jonas Hofmann, Philipp-Florens Lehwalder 等USENIX Security 2026 · 被引用 1 次
- Distributional Private Information RetrievalRyan Lehmkuhl, Alexandra Henzinger, Henry Corrigan-GibbsUSENIX Security 2025
- Found in Translation: A Generative Language Modeling Approach to Memory Access Pattern AttacksGrace Jia, Alex Wong, Anurag KhandelwalUSENIX Security 2025
它引用的顶会 Paper26
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 被引用 1,075 次
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta 等NeurIPS 2021 · 被引用 573 次
- PIR with Compressed Queries and Amortized Query ProcessingSebastian Angel, Hao Chen, Kim Laine, Srinath T. V. SettyS&P 2018 · 被引用 353 次
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas 等MICRO 2021 · 被引用 294 次
- CrypTFlow: Secure TensorFlow InferenceNishant Kumar, Mayank Rathee, Nishanth Chandran, Divya Gupta 等S&P 2020 · 被引用 276 次
相关 Paper
- Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation TrainingYoungeun Kwon, Yunjae Lee, Minsoo RhuHPCA 2021 · 被引用 40 次
- Fleche: an efficient GPU embedding cache for personalized recommendationsMinhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang 等EuroSys 2022 · 被引用 24 次
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam 等MICRO 2024 · 被引用 5 次
- MP-Rec: Hardware-Software Co-design to Enable Multi-path RecommendationSamuel Hsia, Udit Gupta, Bilge Acun, Newsha Ardalani 等ASPLOS 2023 · 被引用 10 次
- RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding ColumnsZaifeng Pan, Zhen Zheng, Feng Zhang, Ruofan Wu 等ASPLOS 2023 · 被引用 7 次
