Lune

EuroSys2026Top-tier venue

TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank Adaptation

Zian Wang, Ziyi Wang, Haonan Jin, Jie Xing, Lanshan Zhang

2026Year

Abstract

With the rapid expansion of large language model inference service users, cloud computing resource costs have become a critical challenge for service providers. Although utilizing end-device resources for auxiliary inference provides new possibilities to reduce cloud computing costs, existing solutions struggle to achieve an ideal balance across multi-task accuracy, end-to-end latency, and cloud computing costs.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get edf54685-6364-42aa-a566-cebed53aaf40

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines