Cost-efficient Collaboration between On-device and Cloud Language Models
Avanika Narayan, Dan Biderman, Sabri Eyuboglu, Avner May, Scott W. Linderman, James Zou, Christopher Ré
Abstract
We investigate an emerging setup in which a small, on-device language model (LM) with access to local data communicates with a frontier, cloudhosted LM to solve real-world tasks involving financial, medical, and scientific reasoning over long documents. Can a local-remote collaboration reduce cloud inference costs while preserving quality? First, we consider a naïve collaboration protocol where the local and remote models simply chat back and forth. Because only the local model ingests the full context, this protocol reduces cloud costs by 30.4×, but recovers only 87% of the performance of the frontier model. We identify two key limitations of this protocol: the local model struggles to (1) follow multiple instructions at once and (2) reason over long contexts. Motivated by these observations, we propose MINIONS, a protocol in which the remote model decomposes the task into easier subtasks over shorter chunks of the document, that are executed locally in parallel. MINIONS reduces costs by 5.7× on average while recovering 97.9% of the remote-only performance. Our analysis reveals several key design choices that influence the tradeoff between cost and performance in local-remote systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bc3545a-0656-4e1a-a1e1-7461147153b6Cited by top-tier papers2
- Boomerang Distillation Enables Zero-Shot Model Size InterpolationSara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero et al.ICLR 2026 · 3 citations
- FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingYujia Fu, Heming Zhong, Dan Huang, Yutong LuACL 2026
Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
Related papers
- Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMsQijun Miao, Zhixuan FangICML 2026
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-TrainingWenzhi Fang, Dong-Jun Han, Liangqi Yuan, Evan Chen et al.ICML 2026 · 4 citations
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device AgentsChenyang Shao, Xinyuan Hu, Yutang Lin, Fengli XuWWW 2025 · 31 citations
- A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language ModelsZuan Xie, Yang Xu, Hongli Xu, Yunming Liao et al.INFOCOM 2026 · 10 citations
- TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank AdaptationZian Wang, Ziyi Wang, Haonan Jin, Jie Xing et al.EuroSys 2026
