Efficient Memory Side-Channel Protection for Embedding Generation in Machine Learning
Muhammad Umar, Akhilesh Parag Marathe, Monami Dutta Gupta, Shubham Jogprakash Ghosh, G. Edward Suh, Wenjie Xiong
摘要
Modern machine learning (ML) models need to process both continuous and categorical/discrete feature values, e.g., deep learning recommendation models (DLRMs) rely on users’ categorical features to make recommendations, and large language models (LLMs) take discrete words/tokens as input. ML models process such discrete features by converting them to numerical vectors called embeddings. Unfortunately, embedding table lookups are vulnerable to side-channel attacks, as table indices leak input feature values. Due to the size of the embedding tables, using conventional oblivious computing techniques such as ORAM to protect memory access patterns to the tables incur significant overhead. In this paper, we propose to use a different technique, Deep Hash Embedding (DHE), to secure embedding table accesses, even though it is not commonly used today due to its compute-intensive nature. We investigate three embedding generation methods with side-channel protection: linear scan of the embedding table, embedding table protected by ORAM, and DHE. Our experiments on DLRMs and LLMs show that DHE or a hybrid scheme combining DHE and linear scan can significantly improve both performance and memory footprint compared to the conventional ORAM protection. For DLRM on Criteo datasets, our hybrid scheme improves performance by about for large embedding tables, and up to end-to-end over the optimized ORAM baseline without any loss in accuracy, while reducing the model memory footprint by up to . For a GPT-2 LLM, using DHE speeds up the prompt prefill by up to and decoding by up to over ORAM, depending on the batch size, with comparable output quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Practical Federated Recommendation Model Learning Using ORAM with Controlled PrivacyJinyu Liu, Wenjie Xiong, G. Edward Suh, Kiwan MaengASPLOS 2025 · 被引用 2 次
- TENNOR: Trustworthy Execution for Neural Networks through Obliviousness and RetrievalsZifan Qu, Vasileios P. Kemerlis, Giuseppe Ateniese, Evgenios M. KornaropoulosCCS 2026
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sanctum: Minimal Hardware Extensions for Strong Software IsolationVictor Costan, Ilia A. Lebedev, Srinivas DevadasUSENIX Security 2016 · 被引用 649 次
- Oblivious Multi-Party Machine Learning on Trusted ProcessorsOlga Ohrimenko, Felix Schuster, Cédric Fournet, Aastha Mehta 等USENIX Security 2016 · 被引用 594 次
- CrypTen: Secure Multi-Party Computation Meets Machine LearningBrian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta 等NeurIPS 2021 · 被引用 573 次
- DRAMA: Exploiting DRAM Addressing for Cross-CPU AttacksPeter Pessl, Daniel Gruss, Clémentine Maurice, Michael Schwarz 等USENIX Security 2016 · 被引用 500 次
相关 Paper
- LAORAM: A Look Ahead ORAM Architecture for Training Large Embedding TablesRachit Rajat, Yongqin Wang, Murali AnnavaramISCA 2023 · 被引用 6 次
- Hybrid Embedding Framework for Memory-Efficient Recommendation SystemsSeung Jin Yang, Hyuk-Jae Lee, Chae-Eun RheeDAC 2025
- Learning to Embed Categorical Features without Embedding Tables for RecommendationWang-Cheng Kang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi 等KDD 2021 · 被引用 46 次
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis 等ASPLOS 2022 · 被引用 65 次
- LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank ObfuscationGaojian Xiong, Yu Sun, Jianhua Liu, Jian Cui 等NeurIPS 2025 · 被引用 6 次
