Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
Shengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou, Xiaowen Chu, Yutong Lu, Xu Chen
摘要
Transformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users’ privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and propose Galaxy, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration. Galaxy introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity-aware parallelism planning for fully exploiting the resource potential. Furthermore, Galaxy devises a tile-based fine-grained overlapping of communication and computation to mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments. Extensive evaluation based on prototype implementation demonstrates that Galaxy remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving up to 2.5× end-to-end latency reduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge DevicesShengyuan Ye, Liekang Zeng, Xiaowen Chu, Guoliang Xing 等MobiCom 2024 · 被引用 29 次
- Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge DevicesShengyuan Ye, Bei Ouyang, Liekang Zeng, Tianyi Qian 等INFOCOM 2025 · 被引用 23 次
- Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache ManagementQianli Liu, Zicong Hong, Peng Li, Fahao Chen 等INFOCOM 2025 · 被引用 4 次
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang 等INFOCOM 2025 · 被引用 3 次
- Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home ClustersZonghang Li, Tao Li, Wenjiao Feng, Rongxing Xiao 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper13
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang 等NeurIPS 2022 · 被引用 345 次
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
相关 Paper
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 被引用 10 次
- HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the EdgeKaiyuan Liu, Lizi Zhang, Chengzhong Xu, Li LiRTSS 2025 · 被引用 1 次
- Mercury: Towards Optimal Accuracy-Latency Trade-off for Collaborative Transformer InferenceYumeng Liang, Jianhui Chang, Sijia Li, Mingyuan Zang 等INFOCOM 2026 · 被引用 1 次
- Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference SystemsAo Zhou, Jianlei Yang, Tong Qiao, Yingjie Qi 等DAC 2024 · 被引用 5 次
- EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge NetworksYiming Yao, Jianwei Niu, Bin Dai, Tao RenACL 2026
