ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
Xiao Liu, Lijun Zhang, Deepak Ganesan, Hui Guan
Abstract
Multi-device inference can reduce Transformer latency by parallelizing computation. However, existing methods require high inter-device bandwidth, making them impractical for bandwidth-constrained environments. We present ASTRA, a communication-efficient framework that integrates sequence parallelism with mixed-precision attention, where non-local token embeddings are transmitted as low-bit vector-quantized codes while local attention remains full precision. To preserve accuracy under aggressive compression, ASTRA introduces Noise-Augmented Quantization and Distributed Class Tokens. Across vision and language models (e.g., ViT and GPT2), ASTRA achieves up to 2.64 speedup over single-device inference and up to 15.25 over prior multi-device baselines while operating at bandwidths as low as 10 Mbps. ASTRA remains robust on large models (e.g., Llama-3-8B) even under non-ideal network conditions such as packet loss and dynamic networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Post-Training Quantization for Vision TransformerZhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang et al.NeurIPS 2021 · 528 citations
- DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Unprecedented ScaleReza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li et al.SC 2022 · 276 citations
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun et al.NeurIPS 2022 · 247 citations
Related papers
- SparQ Attention: Bandwidth-Efficient LLM InferenceLuka Ribar, Ivan Chelombiev, Luke Hudlass-Galley, Charlie Blake et al.ICML 2024 · 108 citations
- Janus: Collaborative Vision Transformer Under Dynamic Network EnvironmentLinyi Jiang, Silvery D. Fu, Yifei Zhu, Bo LiINFOCOM 2025 · 6 citations
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu et al.ASPLOS 2025 · 5 citations
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector QuantizationDingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin et al.ACL 2026 · 4 citations
- Understanding Int4 Quantization for Language Models: Latency Speedup, Composability, and Failure CasesXiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao et al.ICML 2023 · 41 citations
