ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation
Jiangrui Yu, Baosheng Zhang, Liang Kong, Lin Ding, Yi Chen, Ye Yu, Mingzhe Zhang, Meng Li
摘要
Generative large language models (LLMs) have achieved state-of-the-art performance on many real-world tasks such as code generation and question answering. These models predominantly rely on an autoregressive decoding strategy that generates output tokens sequentially. However, their pervasive deployment raises serious privacy concerns, motivating private inference frameworks based on fully homomorphic encryption (FHE). A major limitation of existing FHE frameworks is their inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage. In this paper, we propose ROSETTA, a hybrid CKKS/TFHE framework that overcomes this limitation. We first observe that nonlinear operations in the decode stage exhibit heterogeneous workload patterns, which can be handled effectively via a hybrid approach. We then realize this with two key contributions: 1) an adaptive segmented lookup-table protocol based on TFHE that enables efficient and accurate evaluation of nonlinear operations; and 2) a scheme-aware operator-selection framework that automatically assigns each nonlinear operator to CKKS or TFHE to minimize end-to-end decoding latency. We demonstrate that ROSETTA achieves up to Softmax speedup and -- end-to-end speedup over the SOTA framework CacheMir.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 被引用 1,075 次
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing 等NeurIPS 2022 · 被引用 209 次
- Low-Complexity Deep Convolutional Neural Networks on Fully Homomorphic Encryption Using Multiplexed Parallel ConvolutionsEunsang Lee, Joon-Woo Lee, Junghyun Lee, Young-Sik Kim 等ICML 2022 · 被引用 171 次
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng 等S&P 2024 · 被引用 149 次
相关 Paper
- FxHENN: FPGA-based acceleration framework for homomorphic encrypted CNN inferenceYilan Zhu, Xinyao Wang, Lei Ju, Shanqing GuoHPCA 2023 · 被引用 39 次
- Encryption-Friendly LLM ArchitectureDonghwan Rho, Taeseong Kim, Minje Park, Jung Woo Kim 等ICLR 2025
- EncryptedLLM: Privacy-Preserving Large Language Model Inference via GPU-Accelerated Fully Homomorphic EncryptionLeo de Castro, Daniel Escudero, Adya Agrawal, Antigoni Polychroniadou 等ICML 2025
- Euston: Efficient and User-Friendly Secure Transformer Inference with Non-InteractivityXinwen Gao, Shaojing Fu, Lin Liu, Zhuotao Liu 等S&P 2026 · 被引用 7 次
- SLOTHE : Lazy Approximation of Non-Arithmetic Neural Network Functions over Encrypted DataKevin Nam, Youyeon Joo, Seungjin Ha, Yunheung PaekUSENIX Security 2025
