TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank Adaptation
Zian Wang, Ziyi Wang, Haonan Jin, Jie Xing, Lanshan Zhang
2026Year
Abstract
With the rapid expansion of large language model inference service users, cloud computing resource costs have become a critical challenge for service providers. Although utilizing end-device resources for auxiliary inference provides new possibilities to reduce cloud computing costs, existing solutions struggle to achieve an ideal balance across multi-task accuracy, end-to-end latency, and cloud computing costs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get edf54685-6364-42aa-a566-cebed53aaf40Related papers
- Task-Aware Cloud-End Offloading for Vision-Language Model Serving via Dynamic Modality-Specific Adapter SchedulingZian Wang, Ziyi Wang, Jie Xing, Yaya Wei et al.WWW 2026
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 10 citations
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM InferenceJiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LINICML 2026 · 3 citations
- Cost-efficient Collaboration between On-device and Cloud Language ModelsAvanika Narayan, Dan Biderman, Sabri Eyuboglu, Avner May et al.ICML 2025
