Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models
Chung-ju Huang, Huiqiang Zhao, Yuanpeng He, Lijian Li, Wenpin Jiao, Zhi Jin, Peixuan Chen, Leye Wang
摘要
The increasing reliance on cloud-hosted Large Language Models (LLMs) exposes sensitive client data, such as prompts and responses, to potential privacy breaches by service providers. Existing approaches fail to ensure privacy, maintain model performance, and preserve computational efficiency simultaneously. To address this challenge, we propose Talaria, a confidential inference framework that partitions the LLM pipeline to protect client data without compromising the cloud's model intellectual property or inference quality. Talaria executes sensitive, weight-independent operations within a client-controlled Confidential Virtual Machine (CVM) while offloading weight-dependent computations to the cloud GPUs. The interaction between these environments is secured by our Reversible Masked Outsourcing (ReMO) protocol, which uses a hybrid masking technique to reversibly obscure intermediate data before outsourcing computations. Extensive evaluations show that Talaria can defend against state-of-the-art token inference attacks, reducing token reconstruction accuracy from over 97.5% to an average of 1.34%, all while being a lossless mechanism that guarantees output identical to the original model without significantly decreasing efficiency and scalability. To the best of our knowledge, this is the first work that ensures clients' prompts and responses remain inaccessible to the cloud, while also preserving model privacy, performance, and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt EngineerJunyuan Hong, Jiachen T. Wang, Chenhui Zhang, Zhangheng Li 等ICLR 2024 · 被引用 70 次
- Privacy-Preserving In-Context Learning for Large Language ModelsTong Wu, Ashwinee Panda, Jiachen T. Wang, Prateek MittalICLR 2024 · 被引用 58 次
- PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and RestorationZiqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu 等ACL 2025 · 被引用 31 次
- 00SEVen - Re-enabling Virtual Machine Forensics: Introspecting Confidential VMs Using Privileged in-VM AgentsFabian Schwarz, Christian RossowUSENIX Security 2024 · 被引用 10 次
- CENTAUR: Bridging the Impossible Trinity of Privacy, Efficiency, and Performance in Privacy-Preserving Transformer InferenceJinglong Luo, Guanzhong Chen, Yehong Zhang, Shiyu Liu 等ACL 2025 · 被引用 9 次
相关 Paper
- NOIR: Privacy-Preserving Generation of Code with Open-Source LLMsKhoa Nguyen, Khiem Ton, NhatHai Phan, Issa Khalil 等USENIX Security 2026 · 被引用 2 次
- Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud ServicesZipeng Ye, Wenjian Luo, Qi Zhou, Yubo TangAAAI 2026
- Anti-adversarial Learning: Desensitizing Prompts for Large Language ModelXuan Li, Zhe Yin, Xiaodong Gu, Beijun ShenAAAI 2026
- TextFusion: Privacy-Preserving Pre-trained Model Inference via Token FusionXin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma 等EMNLP 2022 · 被引用 12 次
- Split-and-Denoise: Protect large language model inference with local differential privacyPeihua Mai, Ran Yan, Zhe Huang, Youjia Yang 等ICML 2024 · 被引用 41 次
