Lune

INFOCOM2026顶会

Mercury: Towards Optimal Accuracy-Latency Trade-off for Collaborative Transformer Inference

Yumeng Liang, Jianhui Chang, Sijia Li, Mingyuan Zang, Jie Wu

2026年份
1被引次数

摘要

Vision Transformers (ViT) achieve remarkable performance across a wide range of visual tasks but incur high inference latency on resource-constrained mobile devices. Mobile-cloud collaborative inference, i.e., partially offloading the workload to the cloud, is a promising solution for acceleration by leveraging both mobile and cloud resources. Existing approaches reduce token numbers in a multi-round manner to shrink the data to be offloaded to the cloud, so as to reduce the communication latency. However, they cause imbalanced workloads, leading to heavy latency on the resource-constrained mobile devices. This paper presents Mercury, a collaborative ViT inference framework that accounts for mobile-cloud resource disparity. Unlike multi-round strategies, it performs single-round token pruning at the initial layer on mobile devices to aggressively reduce computation and communication overhead, and then restores tokens in the cloud to maintain accuracy. A similarity-based regionally decentralized token pruning method is designed to preserve semantic features, while an efficient meta-parameter-based token reconstruction method is used to improve accuracy with minimal overhead. Experiments on real-world testbeds under 4G/5G traces show that Mercury achieves up to a 2.27× speedup and halves the mobile-side energy consumption, while preserving comparable accuracy to state-of-the-art methods.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖