QLLMS: Quantization-Adaptive LLM Scheduling for Partially Informed Edge Serving Systems
Miao Hu, Qi He, Di Wu
Abstract
The quantization technologies have enabled the practical deployment of large language models (LLMs) at the edge, however, current edge scheduling and quantization selection are separately designed. Furthermore, partially informed edge processing performance further exacerbates these challenges. To address these issues, we introduce QLLMS, a joint quantization-adaptive scheduling scheme for large language models tailored to partially informed edge serving systems. The primary objective of our approach is to reduce GPU rental costs by strategically orchestrating both quantization options and heterogeneous resources. The QLLMS algorithm first determines the available quantization set (AQS) to optimize the LLM inference performance within limited edge resources. To address the unpredictable nature of edge computing performance, we present a low-rank property-driven recovery approach, which can reconstruct complete AQS matrices using only partial samples. Subsequently, we devise a novel many-to-one matching algorithm that aims to strike a balance between efficient utilization of edge resources and optimal model inference performance, factoring in the available quantization options. We prove that QLLMS yields a stable matching without blocking pairs that could lead to inefficiencies. Experiment results show that QLLMS achieves a reduction of up to 22.36% in rental costs compared to state-of-the-art baselines in partially informed edge serving systems while simultaneously improving the task completion rate on edge servers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Memory-Efficient KV Cache Optimization for Large Language Model Inference at the EdgeChi Zhang, Haisheng Tan, Haotian Pan, Yang Xu et al.INFOCOM 2026 · 1 citation
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui et al.AAAI 2026
Related papers
- Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI ComputingMulei Ma, Chenyu Gong, Liekang Zeng, Yang YangINFOCOM 2025 · 12 citations
- Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUsYouhe Jiang, Fangcheng Fu, Xiaozhe Yao, Guoliang He et al.ICML 2025
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 20 citations
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 10 citations
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language ModelsElias Frantar, Roberto L. Castro, Jiale Chen, Torsten Hoefler et al.PPoPP 2025 · 24 citations
