Janus: Collaborative Vision Transformer Under Dynamic Network Environment
Linyi Jiang, Silvery D. Fu, Yifei Zhu, Bo Li
摘要
Vision Transformers (ViTs) have outperformed traditional Convolutional Neural Network architectures and achieved state-of-the-art results in various computer vision tasks. Since ViTs are computationally expensive, the models either have to be pruned to run on resource-limited edge devices only or have to be executed on remote cloud servers after receiving the raw data transmitted over fluctuating networks. The resulting degraded performance or high latency all hinder their widespread applications. In this paper, we present Janus, the first framework for low-latency cloud-device collaborative Vision Transformer inference over dynamic networks. Janus overcomes the intrinsic model limitations of ViTs and realizes collaboratively executing ViT models on both cloud and edge devices, achieving low latency, high accuracy, and low communication overhead. Specifically, Janus judiciously combines token pruning techniques with a carefully designed fine-to-coarse model splitting policy and non-static mixed pruning policy. It attains a balance between accuracy and latency by dynamically selecting the optimal pruning level and split point. Experimental results across various tasks demonstrate that Janus enhances throughput by up to 5.15x and reduces latency violation ratios by up to 98.7% when compared with baseline approaches under various network environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Hyperion: Low-Latency Ultra-HD Video Analytics via Collaborative Vision Transformer InferenceLinyi Jiang, Yifei Zhu, Hao Yin, Bo LiINFOCOM 2026 · 被引用 1 次
- NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge DevicesZiteng Wei, Qiang He, Bing Li, Feifei Chen 等CVPR 2026 · 被引用 1 次
- Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge IntelligenceZiteng Wei, Qiang He, Feifei Chen, Ranjie Duan 等ICLR 2026
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 被引用 690 次
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision TransformerYifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng 等AAAI 2022 · 被引用 288 次
相关 Paper
- Mercury: Towards Optimal Accuracy-Latency Trade-off for Collaborative Transformer InferenceYumeng Liang, Jianhui Chang, Sijia Li, Mingyuan Zang 等INFOCOM 2026 · 被引用 1 次
- ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-DesignHaoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu 等HPCA 2023 · 被引用 124 次
- HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersPeiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie 等HPCA 2023 · 被引用 117 次
- VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsZhenyu Wang, Hao Luo, Pichao Wang, Feng Ding 等NeurIPS 2022 · 被引用 57 次
- CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision ModelsDenis Kuznedelev, Eldar Kurtic, Elias Frantar, Dan AlistarhNeurIPS 2023 · 被引用 24 次
