SAL-ViT: Towards Latency Efficient Private Inference on ViT using Selective Attention Search with a Learnable Softmax Approximation
Yuke Zhang, Dake Chen, Souvik Kundu, Chenghao Li, Peter A. Beerel
Abstract
Recently, private inference (PI) has addressed the rising concern over data and model privacy in machine learning inference as a service. However, existing PI frameworks suffer from high computational and communication overheads due to the expensive multi-party computation (MPC) protocols, particularly for large models such as vision transformers (ViT). The majority of this overhead is due to the encrypted softmax operation in each self-attention layer. In this work, we present SAL-ViT with two novel techniques to boost PI efficiency on ViTs. Our first technique is a learnable PI-efficient approximation to softmax, namely, learnable 2Quad (L2Q), that introduces learnable scaling and shifting parameters to the prior 2Quad softmax approximation, enabling improvement in accuracy. Then, given our observation that external attention (EA) presents lower PI latency than widely-adopted self-attention (SA) at the cost of accuracy, we present a selective attention search (SAS) method to integrate the strength of EA and SA. Specifically, for a given lightweight EA ViT, we leverage a constrained optimization procedure to selectively search and replace EA modules with SA alternatives to maximize the accuracy. Our extensive experiments show that our SAL-ViT can averagely achieve 1.28×, 1.28×, 1.14× lower PI latency with 1.79%, 1.41%, and 2.08% higher accuracy compared to the existing alternatives, on CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85c42003-e552-46d0-8b68-3a20e5c2827cCited by top-tier papers8
- Ditto: Quantization-aware Secure Inference of Transformers upon MPCHaoqi Wu, Wenjing Fang, Yancheng Zheng, Junming Ma et al.ICML 2024 · 17 citations
- CENTAUR: Bridging the Impossible Trinity of Privacy, Efficiency, and Performance in Privacy-Preserving Transformer InferenceJinglong Luo, Guanzhong Chen, Yehong Zhang, Shiyu Liu et al.ACL 2025 · 9 citations
- MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM InferenceWenxuan Zeng, Ye Dong, Jinjin Zhou, Jin Tan et al.NeurIPS 2025 · 4 citations
- CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-TuningSaisai Xia, Wenhao Wang, Zihao Wang, Yuhui Zhang et al.NDSS 2026 · 4 citations
- Ouromamba: a Data-Free Quantization Framework for Vision MambaAkshat Ramachandran, Mingyu Lee, Huan Xu, Souvik Kundu et al.ICCV 2025 · 3 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
Related papers
- Privacy-Friendly Adaptation of Vision Transformers for Communication and Latency-Efficient Private InferenceZhi Pang, Bo Feng, Meng Luo, Chenhao Liu et al.WWW 2026
- MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous AttentionWenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong et al.ICCV 2023 · 38 citations
- QT-ViT: Improving Linear Attention in ViT with Quadratic Taylor ExpansionYixing Xu, Chao Li, Dong Li, Xiao Sheng et al.NeurIPS 2024 · 7 citations
- MixA: A Mixed Attention Approach with Stable Lightweight Linear Attention to Enhance Efficiency of Vision Transformers at the EdgeSabbir Ahmed, Jingtao Li, Weiming Zhuang, Chen Chen et al.ICCV 2025 · 2 citations
- You Only Need Less Attention at Each Stage in Vision TransformersShuoxi Zhang, Hanpeng Liu, Stephen Lin, Kun HeCVPR 2024 · 19 citations
