Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI Computing
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang
Abstract
Large Language Models (LLMs) have attracted growing attention owing to their advanced capability in under-standing and reacting to instructions. While they are experiencing wide deployment in the multi-tier cloud-edge architecture, their performance is severely constrained by the network capacity, and how to schedule efficient data flow for performance maximum poses significant challenges. Towards that, this paper establishes a comprehensive system model and, for the first time, formulates the efficient LLM scheduling problem in the multi-tier cloud-edge network. Given its non-convexness, we propose the Multi-tier Multi-node Scheduling of LLM (MMSL) algorithm for Collabo-rative AI Computing, a two-stage scheduling framework designed to optimize LLM inference in multi-tier cloud-edge networks. Initially, the inter-tier LLM automated decoupling and partitioning phase employs integer linear programming to allocate model size and computing demands efficiently. Subsequently, the intra-tier LLM task scheduling algorithm, leveraging GNN, identifies optimal scheduling nodes within each tier by evaluating resource utilization and network conditions. Extensive evaluations show that our solution significantly outperforms traditional scheduling methods by 9.1 %- 26.3% throughput improvement.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 801894a6-eb7b-42a7-b8bc-c4b6f17a463eRelated papers
- QLLMS: Quantization-Adaptive LLM Scheduling for Partially Informed Edge Serving SystemsMiao Hu, Qi He, Di WuINFOCOM 2025 · 8 citations
- HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the EdgeKaiyuan Liu, Lizi Zhang, Chengzhong Xu, Li LiRTSS 2025 · 1 citation
- DEMUS: Large Multimodal Model Serving at the Edge with Diffusion-Based SchedulingHan Zhang, Xiangkai Ma, Tiantian Wang, Mingkai Lin et al.INFOCOM 2026
- Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUsYouhe Jiang, Fangcheng Fu, Xiaozhe Yao, Guoliang He et al.ICML 2025
- Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge DevicesShengyuan Ye, Bei Ouyang, Liekang Zeng, Tianyi Qian et al.INFOCOM 2025 · 23 citations
