Tricycle: Private Transformer Inference with Tricyclic Encodings
Lawrence Lim, Vikas Kalagi, Julia Novick, Jiaming Liu, Divyakant Agrawal, Amr El Abbadi
Abstract
The growing deployment of large language models (LLMs) in privacy-sensitive settings demands inference mechanisms that preserve data confidentiality. Homomorphic encryption (HE) offers a principled solution by enabling computation directly on encrypted data; however, existing solutions struggle to scale to full LLMs. We present Tricycle, a system for efficient private transformer inference. At its core, Tricycle introduces tricyclic encodings, a novel packing scheme that enables batch matrix multiplications with optimal multiplicative depth while naturally supporting multi-head attention. We demonstrate that Tricycle's matrix multiplications are particularly amenable for integration with other optimizations including Baby-Step Giant-Step optimizations, optimized block matrix multiplications, lazy relinearization, and complexification, which significantly improve performance. We further introduce statistical max estimation, a lightweight method for stabilizing softmax under HE. We implement Tricycle end-to-end on a GPU-accelerated CKKS pipeline and evaluate it on BERT models. For BERT-Base with 128 tokens, Tricycle achieves 100.5 seconds latency on a single GPU, yielding and speedups over prior state-of-the-art systems, Thor and Powerformer, respectively. These results demonstrate that careful design of packing, algorithms, and optimizations can significantly reduce the cost of private LLM inference, bringing practical deployment closer to reality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c962c318-01ac-4639-bb4c-d4970c60bf07Cited by top-tier papers3
- MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer InferenceLinru Zhang, Xiangning Wang, Sim Jun Jie, Zhicong Huang et al.ICLR 2026 · 24 citations
- Hyperion: Private Token Sampling with Homomorphic EncryptionLawrence Lim, Jiaming Liu, Vikas Kalagi, Divyakant Agrawal et al.ACL 2026 · 1 citation
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang et al.USENIX Security 2026
Related papers
- THOR: Secure Transformer Inference with Homomorphic EncryptionJungho Moon, Dongwoo Yoo, Xiaoqian Jiang, Miran KimCCS 2025 · 1 citation
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCTianshi Xu, Wen-jie Lu, Jiangrui Yu, Yi Chen et al.USENIX Security 2025
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
- Primer: Fast Private Transformer Inference on Encrypted DataMengxin Zheng, Qian Lou, Lei JiangDAC 2023 · 26 citations
- Powerformer: Efficient and High-Accuracy Privacy-Preserving Language Model with Homomorphic EncryptionDongjin Park, Eunsang Lee, Joon-Woo LeeACL 2025 · 13 citations
