CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
Chunlin Tian, Xinpeng Qin, Kahou Tam, Li Li, Zijian Wang, Yuanzhe Zhao, Minglei Zhang, Chengzhong Xu
Abstract
Deploying large language models (LLMs) on edge devices is crucial for delivering fast responses and ensuring data privacy. However, the limited storage, weight, and power of edge devices make it difficult to deploy LLM-powered applications. These devices must balance latency requirements with energy consumption and model accuracy. In this paper, we first quantify the challenges of deploying LLMs on off-the-shelf edge devices and then we present CLONE, an in-depth algorithm-hardware co-design at both the model- and system-level that intelligently integrates real-time, energy optimization while maintaining robust generality. In order to maximize the synergistic benefits of these algorithms in always-on and intermediate edge computing settings, we specialize in a 28nm scalable hardware accelerator system. We implement and extensively evaluate CLONE on two off-the-shelf edge platforms. Experiments show that CLONE effectively accelerates the inference process up to 11.92x, and saves energy up to 7.36x, while maintaining high-generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45157a88-cfe7-41ca-9cfe-226397b51e43Cited by top-tier papers3
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM InferenceJiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LINICML 2026 · 3 citations
- Streaming Video Crime Anticipation with Spatio-Temporal Causal ReasoningYusong Wang, Zheyuan Gu, Keyu Mao, Minghao Shao et al.CVPR 2026 · 1 citation
- FARLock: Asymmetric RDMA Locking Made FairYuehao Hu, Jiatang Zhou, Tianzheng Wang, Keval VoraOSDI 2026
Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 2 citations
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 10 citations
- CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge DeploymentQinfeng Li, Tianyue Luo, Xuhong Zhang, Yangfan Xie et al.NeurIPS 2025 · 9 citations
- VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible AcceleratorZhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang et al.DAC 2025 · 1 citation
- A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge SystemsYuze Liu, Yunhan Wang, Tiehua Zhang, Zhishu Shen et al.WWW 2026 · 1 citation
