DyMerge-LoRA: On-GPU Post-Merge Fusion for High-Throughput Multi-Tenant Composite LoRA Serving
Rui Xu, Long Chen, Huazheng Lao, Jinquan Zhang, Xia Zhu
摘要
In this paper, we consider the problem of efficiently serving multi-tenant requests for large language models (LLMs) equipped with compositions of multiple Low-Rank Adaptation (LoRA) adapters, aiming to optimize inference throughput, latency, and GPU memory usage. Existing multi-tenant inference systems are primarily architected for a one-to-one mapping between requests and single adapters, limiting their efficiency in handling compositions of multiple LoRA adapters. To address these issues, we propose DyMerge-LoRA, a novel inference serving system specifically designed for high-throughput multi-tenant scenarios by dynamically fusing corresponding LoRA adapters on the GPU. DyMerge-LoRA introduces a Base Adapter Manager (BAM) for efficient GPU memory management of individual basic adapters, a LoRA-Aware Greedy Scheduler (LAG) for batching, and a Mapping-based Adapter Fusion Matrix-Vector Multiplication (MAF-MVM) kernel for high computational throughput. Experimental evaluations conducted on multiple setups demonstrate that DyMerge-LoRA achieves up to 1.1-4× improvement in inference throughput compared to state-of-the-art methods, significantly reduces TTFT, and effectively mitigates GPU memory consumption, validating its effectiveness and performance advantages under diverse workload conditions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun 等OSDI 2024 · 被引用 79 次
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen 等HPCA 2026
- Compress then Serve: Serving Thousands of LoRA Adapters with Little OverheadRickard Brüel Gabrielsson, Jiacheng Zhu, Onkar Bhardwaj, Leshem Choshen 等ICML 2025
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and ServingYuchen Zhang, Hanyue Du, Chun Cao, Jingwei XuNeurIPS 2025 · 被引用 1 次
- Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM InferenceSuyi Li, Hanfeng Lu, Tianyuan Wu, Minchen Yu 等USENIX ATC 2025 · 被引用 22 次
