Mercury: Towards Optimal Accuracy-Latency Trade-off for Collaborative Transformer Inference
Yumeng Liang, Jianhui Chang, Sijia Li, Mingyuan Zang, Jie Wu
Abstract
Vision Transformers (ViT) achieve remarkable performance across a wide range of visual tasks but incur high inference latency on resource-constrained mobile devices. Mobile-cloud collaborative inference, i.e., partially offloading the workload to the cloud, is a promising solution for acceleration by leveraging both mobile and cloud resources. Existing approaches reduce token numbers in a multi-round manner to shrink the data to be offloaded to the cloud, so as to reduce the communication latency. However, they cause imbalanced workloads, leading to heavy latency on the resource-constrained mobile devices. This paper presents Mercury, a collaborative ViT inference framework that accounts for mobile-cloud resource disparity. Unlike multi-round strategies, it performs single-round token pruning at the initial layer on mobile devices to aggressively reduce computation and communication overhead, and then restores tokens in the cloud to maintain accuracy. A similarity-based regionally decentralized token pruning method is designed to preserve semantic features, while an efficient meta-parameter-based token reconstruction method is used to improve accuracy with minimal overhead. Experiments on real-world testbeds under 4G/5G traces show that Mercury achieves up to a 2.27× speedup and halves the mobile-side energy consumption, while preserving comparable accuracy to state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f2737c2e-7017-4cf3-b558-5a8a1a728457Related papers
- Janus: Collaborative Vision Transformer Under Dynamic Network EnvironmentLinyi Jiang, Silvery D. Fu, Yifei Zhu, Bo LiINFOCOM 2025 · 6 citations
- Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile DevicesZhihao Shu, Xiaowei Yu, Zihao Wu, Wenqi Jia et al.NeurIPS 2024 · 6 citations
- HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersPeiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie et al.HPCA 2023 · 117 citations
- ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative PruningWen Luo, Peng Chen, Xiaotao Huang, LiQun HuangAAAI 2026
- V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision TransformerGuangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang et al.AAAI 2026 · 1 citation
