AIMS: Cost-Efficient LLM-Based Agent Deployment in Hybrid Cloud-Edge Environments
Shiyi Liu, Haiying Shen, Shuai Che, Mahdi Ghandi, Mingqin Li
摘要
In the realm of AI, large language models (LLMs) like GPT-5, central to the operation of AI agents, predominantly operate in the cloud, incurring high operational costs. With local-based small language models (SLMs) becoming more accurate, the necessity of cloud-exclusive processing is being reconsidered. An AI agent's response to a user's request comprises a series of subtasks or iterations. Existing approaches choose either an LLM or SLM for the entire request to ensure similar outputs, but this is ineffective for AI agents as SLMs may generate differing subtasks, compromising final accuracy. In this paper, we first conduct experimental analysis to understand the features of AI agent operations. Leveraging our findings, we propose the Adaptive Iteration-level Model Selector (AIMS), a lightweight scheduler to automatically partition an AI agent's subtasks between local-based SLM and cloud-based LLM. AIMS considers the varying subtask features and strategically decides the location for each subtask in order to use SLM as much as possible while maintaining the accuracy level. Our experimental results demonstrate that AIMS achieves up to a 27.5% relative improvement in accuracy and up to 31.4% relative increase in SLM usage compared to HybridLLM. It offloads 83.4% of subtasks to a local SLM while attaining similar accuracy on average compared with the cloud-only LLM approach.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative LearningHao Sun, Jiayi Wu, Hengyi Cai, Xiaochi Wei 等EMNLP 2024
- Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMsQijun Miao, Zhixuan FangICML 2026
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMsGucongcong Fan, Chaoyue Niu, Chengfei Lyu, Fan Wu 等NeurIPS 2025 · 被引用 9 次
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device AgentsChenyang Shao, Xinyuan Hu, Yutang Lin, Fengli XuWWW 2025 · 被引用 31 次
- Transcending Cost-Quality Tradeoff in Agent Serving via Session-AwarenessYanyu Ren, Li Chen, Dan Li, Xizheng Wang 等NeurIPS 2025 · 被引用 5 次
