Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents
Chenyang Shao, Xinyuan Hu, Yutang Lin, Fengli Xu
Abstract
The rapid expansion of web content has made on-device AI assistants indispensable for helping users manage the increasing complexity of online tasks. The emergent reasoning ability in large language models offer a promising path for next-generation on-device AI agents. However, deploying full-scale Large Language Models (LLMs) on resource-limited local devices is challenging. In this paper, we propose Division-o f-Thoughts (DoT), a collaborative reasoning framework leveraging the synergy between locally deployed Smaller-scale Language Models (SLMs) and cloud-based LLMs. DoT leverages a Task Decomposer to elicit the inherent planning abilities in language models to decompose user queries into smaller sub-tasks, which allows hybrid language models to fully exploit their respective strengths. Besides, DoT employs a Task Scheduler to analyze the pair-wise dependency of sub-tasks and create a dependency graph, facilitating parallel reasoning of sub-tasks and the identification of key steps. To allocate the appropriate model based on the difficulty of sub-tasks, DoT leverages a Plug-and-Play Adapter, which is an additional task head attached to the SLM that does not alter the SLM's parameters. To boost adapter's task allocation capability, we propose a self-reinforced training method that relies solely on task execution feedback. Extensive experiments on various benchmarks demonstrate that our DoT significantly reduces LLM costs while maintaining competitive reasoning accuracy. Specifically, DoT reduces the average reasoning time and API costs by 66.12% and 83.57%, while achieving comparable reasoning accuracy with the best baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1115fb14-8cce-4541-8bbd-ef74a419c3c3Cited by top-tier papers7
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-TrainingWenzhi Fang, Dong-Jun Han, Liangqi Yuan, Evan Chen et al.ICML 2026 · 4 citations
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM InferenceJiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LINICML 2026 · 3 citations
- Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language ModelsPeijie Liu, Fengli Xu, Yong LiICML 2025
- Adaptive Model and Strategy Routing for Cost-Efficient LLM ServicesZhihong Pan, Kai Zhang, Yuze Zhao, Yupeng HanWWW 2026
- AgentSquare: Automatic LLM Agent Search in Modular Design SpaceYu Shang, Yu Li, Keyu Zhao, Likai Ma et al.ICLR 2025
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
Related papers
- AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative LearningHao Sun, Jiayi Wu, Hengyi Cai, Xiaochi Wei et al.EMNLP 2024
- DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories SearchMurong Yue, Wenlin Yao, Haitao Mi, Dian Yu et al.ICLR 2025
- Cost-efficient Collaboration between On-device and Cloud Language ModelsAvanika Narayan, Dan Biderman, Sabri Eyuboglu, Avner May et al.ICML 2025
- Latent-Guided Reasoning: Empowering Small LLMs with Large-Model ThinkingHanzhu Chen, Lin Yang, Jie Wang, Junhao Yan et al.ICLR 2026
- AIMS: Cost-Efficient LLM-Based Agent Deployment in Hybrid Cloud-Edge EnvironmentsShiyi Liu, Haiying Shen, Shuai Che, Mahdi Ghandi et al.EuroSys 2026 · 1 citation
