Lune

KDD2026顶会

DyMerge-LoRA: On-GPU Post-Merge Fusion for High-Throughput Multi-Tenant Composite LoRA Serving

Rui Xu, Long Chen, Huazheng Lao, Jinquan Zhang, Xia Zhu

2026年份

摘要

In this paper, we consider the problem of efficiently serving multi-tenant requests for large language models (LLMs) equipped with compositions of multiple Low-Rank Adaptation (LoRA) adapters, aiming to optimize inference throughput, latency, and GPU memory usage. Existing multi-tenant inference systems are primarily architected for a one-to-one mapping between requests and single adapters, limiting their efficiency in handling compositions of multiple LoRA adapters. To address these issues, we propose DyMerge-LoRA, a novel inference serving system specifically designed for high-throughput multi-tenant scenarios by dynamically fusing corresponding LoRA adapters on the GPU. DyMerge-LoRA introduces a Base Adapter Manager (BAM) for efficient GPU memory management of individual basic adapters, a LoRA-Aware Greedy Scheduler (LAG) for batching, and a Mapping-based Adapter Fusion Matrix-Vector Multiplication (MAF-MVM) kernel for high computational throughput. Experimental evaluations conducted on multiple setups demonstrate that DyMerge-LoRA achieves up to 1.1-4× improvement in inference throughput compared to state-of-the-art methods, significantly reduces TTFT, and effectively mitigates GPU memory consumption, validating its effectiveness and performance advantages under diverse workload conditions.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖