QLLMS: Quantization-Adaptive LLM Scheduling for Partially Informed Edge Serving Systems
Miao Hu, Qi He, Di Wu
摘要
The quantization technologies have enabled the practical deployment of large language models (LLMs) at the edge, however, current edge scheduling and quantization selection are separately designed. Furthermore, partially informed edge processing performance further exacerbates these challenges. To address these issues, we introduce QLLMS, a joint quantization-adaptive scheduling scheme for large language models tailored to partially informed edge serving systems. The primary objective of our approach is to reduce GPU rental costs by strategically orchestrating both quantization options and heterogeneous resources. The QLLMS algorithm first determines the available quantization set (AQS) to optimize the LLM inference performance within limited edge resources. To address the unpredictable nature of edge computing performance, we present a low-rank property-driven recovery approach, which can reconstruct complete AQS matrices using only partial samples. Subsequently, we devise a novel many-to-one matching algorithm that aims to strike a balance between efficient utilization of edge resources and optimal model inference performance, factoring in the available quantization options. We prove that QLLMS yields a stable matching without blocking pairs that could lead to inefficiencies. Experiment results show that QLLMS achieves a reduction of up to 22.36% in rental costs compared to state-of-the-art baselines in partially informed edge serving systems while simultaneously improving the task completion rate on edge servers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Memory-Efficient KV Cache Optimization for Large Language Model Inference at the EdgeChi Zhang, Haisheng Tan, Haotian Pan, Yang Xu 等INFOCOM 2026 · 被引用 1 次
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui 等AAAI 2026
相关 Paper
- Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI ComputingMulei Ma, Chenyu Gong, Liekang Zeng, Yang YangINFOCOM 2025 · 被引用 12 次
- Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUsYouhe Jiang, Fangcheng Fu, Xiaozhe Yao, Guoliang He 等ICML 2025
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 被引用 10 次
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language ModelsElias Frantar, Roberto L. Castro, Jiale Chen, Torsten Hoefler 等PPoPP 2025 · 被引用 24 次
