S andhi : Fine-Grained Merging for Memory Efficient Multi-Model Serving
Vima Gupta, Oytun Kuday Duran, Nandan Suresh Meda, Ikhyun An, Ganesh Ananthanarayanan, Anand Iyer
摘要
As autoregressive models become adept at handling various domain-specific tasks, it is necessary to deploy several fine-tuned models concurrently in real-world scenarios. Such deployments are limited by GPU memory, which dictates the cost (i.e., how many models can be hosted) and their performance (i.e., latency and throughput). In this paper, we propose model merging as a way to reduce the memory footprint of co-located models. While model merging in a traditional sense—where multiple models are combined to create a single model to instill emergent behaviors—may seem like a natural fit for memory reduction, we show that directly extending it results in unacceptable accuracy drops. We present Sandhi, a system that adaptively merges models, at a component granularity, while adhering to user's accuracy requirements. Our evaluation on 12 models spanning 3 model families across 9 different benchmarks shows that Sandhi reduces GPU memory footprint by up to 49.8%, which translates to improvements of up to 2.93× in throughput and 2× lower cost.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the EdgeArthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan 等NSDI 2023 · 被引用 94 次
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 被引用 35 次
- DyMerge-LoRA: On-GPU Post-Merge Fusion for High-Throughput Multi-Tenant Composite LoRA ServingRui Xu, Long Chen, Huazheng Lao, Jinquan Zhang 等KDD 2026
- KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference ServersMarcus Chow, Ali Jahanshahi, Daniel WongHPCA 2023 · 被引用 25 次
- Free-Merging: Fourier Transform for Efficient Model MergingShenghe Zheng, Hongzhi WangICCV 2025 · 被引用 12 次
