EcoTune: Edge-Cloud Collaborative Model Adaptation for Budget-Constrained On-Device SLM Personalization
Gong Chen, Mingkai Lin, Xiaobin Hong, Wenzhong Li, Sanglu Lu
Abstract
The rapid growth of web content has spurred the widespread adoption of on-device AI assistants powered by large language models (LLMs). However, deploying and personalizing these assistants in real-world environments remains challenging due to limited annotation budgets and scarce on-device fine-tuning resources. Existing edge–cloud collaboration frameworks typically rely on costly cloud-based supervision or perform full-layer finetuning, leading to inefficiencies in both computation and adaptation. To address these limitations, we propose EcoTune, a budget-constrained framework for efficient edge–cloud collaborative adaptation. EcoTune jointly optimizes representative data selection for cloud annotation and selective on-device model adaptation within a unified closed-loop process. Specifically, it employs a multi-armed bandit–based strategy to identify highvalue user interactions for cloud supervision and a layer importance–driven adaptation mechanism to update only critical components of the small language model (SLM). This coordinated optimization enables dynamic, resource-efficient personalization under stringent annotation and tuning budgets. Experiments on real-world testbeds demonstrate that EcoTune achieves up to 20%-60% reduction in annotation costs and significantly lowers fine-tuning memory consumption compared to state-of-the-art baselines, providing a practical and scalable solution for personalized on-device LLMs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3e576c65-0abc-41e8-b8ee-924ea1717895Related papers
- Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and SynthesisRuiyang Qin, Jun Xia, Zhenge Jia, Meng Jiang et al.DAC 2024 · 17 citations
- LaTune: Lightweight and Adaptive Configuration Tuning for LLM Inference on Edge DevicesSiqi Zhong, Mugeng Liu, Haiyang Shen, Chongyang Pan et al.WWW 2026
- FedDynMask: Efficient Federated Fine-Tuning for Edge LLMs via Dynamic Sparse MaskingYan Wang, Ziyi Gao, Yida Zhang, Rui WangINFOCOM 2026
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Unified Compression and Adaptive Layer VotingZhongzhi Yu, Zheng Wang, Yuhan Li, Ruijie Gao et al.DAC 2024 · 57 citations
- MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference EnginesLei Gao, Amir Ziashahabi, Yue Niu, Salman Avestimehr et al.EMNLP 2025 · 1 citation
