AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning
Hao Sun, Jiayi Wu, Hengyi Cai, Xiaochi Wei, Yue Feng, Bo Wang, Shuaiqiang Wang, Yan Zhang, Dawei Yin
摘要
Recent advancements in large language models (LLMs) have been remarkable. Users face a choice between using cloud-based LLMs for generation quality and deploying local-based LLMs for lower computational cost. The former option is typically costly and inefficient, while the latter usually fails to deliver satisfactory performance for reasoning steps requiring deliberate thought processes. In this work, we propose a novel LLM utilization paradigm that facilitates the collaborative operation of large cloud-based LLMs and smaller local-deployed LLMs. Our framework comprises two primary modules: the local agent instantiated with a relatively smaller LLM, handling less complex reasoning steps, and the cloud agent equipped with a larger LLM, managing more intricate reasoning steps. This collaborative processing is enabled through an adaptive mechanism where the local agent introspectively identifies errors and proactively seeks assistance from the cloud agent, thereby effectively integrating the strengths of both locally-deployed and cloudbased LLMs, resulting in significant enhancements in task completion performance and efficiency. We evaluate ADASWITCH across 7 benchmarks, ranging from mathematical reasoning and complex question answering, using various types of LLMs to instantiate the local and cloud agents. The empirical results show that ADASWITCH effectively improves the performance of the local agent, and sometimes achieves competitive results compared to the cloud agent while utilizing much less computational overhead. Question: Joy is 2 years older than twice the age of Tom. If Tom is 10 years old, how old is Joy? Thought: Calculate twice the age of Tom.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Reasoning Can Be Restored by Correcting a Few Decision TokensShen Changshuo, Leheng Sheng, Yuxin Chen, Xiang Wang 等ICML 2026 · 被引用 2 次
- Efficient Thought Space Exploration Through Strategic InterventionZiheng Li, Hengyi Cai, Xiaochi Wei, Yuchen Li 等AAAI 2026 · 被引用 1 次
- AlphaRouter: Token-level Routing Between SLM and LLM with Reinforcement Learning and Tree SearchSiteng Liao, Yuzhu Liang, Hengzhong Rao, Xizhao Luo 等ICML 2026
它引用的顶会 Paper12
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksBill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman 等NeurIPS 2023 · 被引用 244 次
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju 等NeurIPS 2024 · 被引用 145 次
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
相关 Paper
- Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMsQijun Miao, Zhixuan FangICML 2026
- AIMS: Cost-Efficient LLM-Based Agent Deployment in Hybrid Cloud-Edge EnvironmentsShiyi Liu, Haiying Shen, Shuai Che, Mahdi Ghandi 等EuroSys 2026 · 被引用 1 次
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device AgentsChenyang Shao, Xinyuan Hu, Yutang Lin, Fengli XuWWW 2025 · 被引用 31 次
- Adadrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous DrivingRuifei Zhang, Junlin Xie, Wei Zhang, Weikai Chen 等ICCV 2025 · 被引用 3 次
- Latent-Guided Reasoning: Empowering Small LLMs with Large-Model ThinkingHanzhu Chen, Lin Yang, Jie Wang, Junhao Yan 等ICLR 2026
