Enhancing On-Device LLM Inference with Historical Cloud-Based LLM Interactions
Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, Guihai Chen
摘要
Many billion-scale large language models (LLMs) have been released for resource-constraint mobile devices to provide local LLM inference service when cloud-based powerful LLMs are not available. However, the capabilities of current on-device LLMs still lag behind those of cloud-based LLMs, and how to effectively and efficiently enhance on-device LLM inference becomes a practical requirement. We thus propose to collect the user's historical interactions with the cloud-based LLM and build an external datastore on the mobile device for enhancement using nearest neighbors search. Nevertheless, the full datastore improves the quality of token generation at the unacceptable expense of much slower generation speed. To balance performance and efficiency, we propose to select an optimal subset of the full datastore within the given size limit, the optimization objective of which is proven to be submodular. We further design an offline algorithm, which selects the subset after the construction of the full datastore, as well as an online algorithm, which performs selection over the stream and can be flexibly scheduled. We theoretically analyze the performance guarantee and the time complexity of the offline and the online designs to demonstrate effectiveness and scalability. We finally take three ChatGPT related dialogue datasets and four different on-device LLMs for evaluation. Evaluation results show that the proposed designs significantly enhance LLM performance in terms of perplexity while maintaining fast token generation speed. Practical overhead testing on the smartphone reveal the efficiency of on-device datastore subset selection from memory usage and computation overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge ComputingGuihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou 等RTSS 2025 · 被引用 4 次
- LSRP: A Leader-Subordinate Retrieval Framework for Privacy-Preserving Cloud-Device CollaborationYingyi Zhang, Pengyue Jia, Xianneng Li, Derong Xu 等KDD 2025 · 被引用 2 次
- Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge IntelligenceZiteng Wei, Qiang He, Feifei Chen, Ranjie Duan 等ICLR 2026
- On-Device Collaborative Language Modeling via a Mixture of Generalists and SpecialistsDongyang Fan, Bettina Messmer, Nikita Doikov, Martin JaggiICML 2025
- AGRAG: Advanced Graph-Based Retrieval-Augmented Generation for LLMsYubo Wang, Haoyang Li, Fei Teng, Lei ChenICDE 2026
相关 Paper
- Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid InferenceJihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang 等ACL 2024 · 被引用 2 次
- SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?Xudong Lu, Haohao Gao, Renshou Wu, Shuai Ren 等EMNLP 2025
- Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and SynthesisRuiyang Qin, Jun Xia, Zhenge Jia, Meng Jiang 等DAC 2024 · 被引用 17 次
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-TrainingWenzhi Fang, Dong-Jun Han, Liangqi Yuan, Evan Chen 等ICML 2026 · 被引用 4 次
- Scaling Retrieval-Based Language Models with a Trillion-Token DatastoreRulin Shao, Jacqueline He, Akari Asai, Weijia Shi 等NeurIPS 2024 · 被引用 76 次
