LaTune: Lightweight and Adaptive Configuration Tuning for LLM Inference on Edge Devices
Siqi Zhong, Mugeng Liu, Haiyang Shen, Chongyang Pan, Yun Ma
Abstract
Large Language Models (LLMs) are increasingly deployed on edge devices to address privacy and latency concerns in modern Web applications. While numerous studies focus on inference frameworks, the critical problem of tuning runtime configurations remains largely underexplored. This endeavor is particularly challenging on edge devices due to severe budget limitations and the dynamic variability of system resources. To address these challenges, we draw upon key insights regarding parameter sensitivity, configuration transferability, and rank stability to propose LaTune, a lightweight and adaptive tuning framework. LaTune is designed to efficiently find optimal runtime configurations by incorporating three complementary components: parameter selection to focus on the most impactful parameters, knowledge transfer to leverage historical data for accelerated search, and two-stage optimization to dynamically select the best configuration based on real-time resource constraints. Experiments across four edge devices and LLMs show that LaTune achieves up to 3.93x higher hypervolume and 6.90x throughput gains over baselines. It accelerates tuning efficiency by 2-3x, converging within 10-20 iterations, and ensures robust execution under heavy contention where static methods fail. Our code is open-sourced at https://github.com/pkuaiweb/LaTune.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 01aa86cf-fb84-4ba9-92d8-e4bdc8246a46Related papers
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Unified Compression and Adaptive Layer VotingZhongzhi Yu, Zheng Wang, Yuhan Li, Ruijie Gao et al.DAC 2024 · 57 citations
- EcoTune: Edge-Cloud Collaborative Model Adaptation for Budget-Constrained On-Device SLM PersonalizationGong Chen, Mingkai Lin, Xiaobin Hong, Wenzhong Li et al.WWW 2026
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li et al.USENIX ATC 2025 · 41 citations
- AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the EdgeTao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou et al.INFOCOM 2025 · 5 citations
- HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the EdgeKaiyuan Liu, Lizi Zhang, Chengzhong Xu, Li LiRTSS 2025 · 1 citation
