Lune

SOSP2026Top-tier venue

S andhi : Fine-Grained Merging for Memory Efficient Multi-Model Serving

Vima Gupta, Oytun Kuday Duran, Nandan Suresh Meda, Ikhyun An, Ganesh Ananthanarayanan, Anand Iyer

2026Year

Abstract

As autoregressive models become adept at handling various domain-specific tasks, it is necessary to deploy several fine-tuned models concurrently in real-world scenarios. Such deployments are limited by GPU memory, which dictates the cost (i.e., how many models can be hosted) and their performance (i.e., latency and throughput). In this paper, we propose model merging as a way to reduce the memory footprint of co-located models. While model merging in a traditional sense—where multiple models are combined to create a single model to instill emergent behaviors—may seem like a natural fit for memory reduction, we show that directly extending it results in unacceptable accuracy drops. We present Sandhi, a system that adaptively merges models, at a component granularity, while adhering to user's accuracy requirements. Our evaluation on 12 models spanning 3 model families across 9 different benchmarks shows that Sandhi reduces GPU memory footprint by up to 49.8%, which translates to improvements of up to 2.93× in throughput and 2× lower cost.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines