DyMerge-LoRA: On-GPU Post-Merge Fusion for High-Throughput Multi-Tenant Composite LoRA Serving
Rui Xu, Long Chen, Huazheng Lao, Jinquan Zhang, Xia Zhu
Abstract
In this paper, we consider the problem of efficiently serving multi-tenant requests for large language models (LLMs) equipped with compositions of multiple Low-Rank Adaptation (LoRA) adapters, aiming to optimize inference throughput, latency, and GPU memory usage. Existing multi-tenant inference systems are primarily architected for a one-to-one mapping between requests and single adapters, limiting their efficiency in handling compositions of multiple LoRA adapters. To address these issues, we propose DyMerge-LoRA, a novel inference serving system specifically designed for high-throughput multi-tenant scenarios by dynamically fusing corresponding LoRA adapters on the GPU. DyMerge-LoRA introduces a Base Adapter Manager (BAM) for efficient GPU memory management of individual basic adapters, a LoRA-Aware Greedy Scheduler (LAG) for batching, and a Mapping-based Adapter Fusion Matrix-Vector Multiplication (MAF-MVM) kernel for high computational throughput. Experimental evaluations conducted on multiple setups demonstrate that DyMerge-LoRA achieves up to 1.1-4× improvement in inference throughput compared to state-of-the-art methods, significantly reduces TTFT, and effectively mitigates GPU memory consumption, validating its effectiveness and performance advantages under diverse workload conditions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun et al.OSDI 2024 · 79 citations
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen et al.HPCA 2026
- Compress then Serve: Serving Thousands of LoRA Adapters with Little OverheadRickard Brüel Gabrielsson, Jiacheng Zhu, Onkar Bhardwaj, Leshem Choshen et al.ICML 2025
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and ServingYuchen Zhang, Hanyue Du, Chun Cao, Jingwei XuNeurIPS 2025 · 1 citation
- Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM InferenceSuyi Li, Hanfeng Lu, Tianyuan Wu, Minchen Yu et al.USENIX ATC 2025 · 22 citations
