MOSAIC: Masked Outsourcing of Secure AI Computations
James Hsin-yu Chiang, Sheila Zingg, Kari Kostiainen, Srdjan Capkun
摘要
We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server must learn neither. We present MOSAIC, whose core is a novel matrix-multiplication masking protocol that scales to far larger matrices than prior work, enabling the safe outsourcing of modern workloads such as large transformer inference. By introducing small amounts of noise to the multiplication result and thereby relaxing correctness, MOSAIC achieves optimal asymptotic client overhead and concrete runtimes orders of magnitude faster than prior work. Its security reduces to the decisional LWE and LPN assumptions.
Because this noise accumulates across the many layers of a transformer, a key technical challenge is bounding error growth; MO-SAIC addresses this with an error-scaling mechanism based on random Hadamard rotations. On large 70B transformer models, MOSAIC's perplexity is comparable to popular quantization approaches and even matches full-precision BF16 inference on Hu-manEval.
Finally, we present an end-to-end implementation 1 showing how ideas like MOSAIC can promise a path towards large-scale confidential AI in modern data centers. Non-confidential inference is already distributed across phase (prefill/decode), layer, and time to maximize utilization of heterogeneous hardware, using RDMAlike networking to move activations, cached KV values, and weights across nodes. MOSAIC enables scaling of confidential compute by keeping the trusted computing base (TCB) small and outsourcing the bulk of the AI computation to untrusted accelerators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 被引用 1,075 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- Secure Outsourced Matrix Computation and Application to Neural NetworksXiaoqian Jiang, Miran Kim, Kristin E. Lauter, Yongsoo SongCCS 2018 · 被引用 359 次
- SOTER: Guarding Black-box Inference for General Neural Networks at the EdgeTianxiang Shen, Ji Qi, Jianyu Jiang, Xian Wang 等USENIX ATC 2022 · 被引用 67 次
相关 Paper
- SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-ComputeBowen Shen, Yuyue Chen, Peng Yang, Bin Zhang 等AAAI 2026
- Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language ModelsChung-ju Huang, Huiqiang Zhao, Yuanpeng He, Lijian Li 等ACL 2026
- Secure Outlier-Aware Large Language Model InferenceLifan Zhao, Zhixuan FangICLR 2026
- BOLT: Privacy-Preserving, Accurate and Efficient Inference for TransformersQi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng 等S&P 2024 · 被引用 149 次
- Mosformer: Maliciously Secure Three-Party Inference Framework for Large TransformersKe Cheng, Yuheng Xia, Anxiao Song, Jiaxuan Fu 等CCS 2025
