USENIX ATC2025顶会
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
Chunlin Tian, Xinpeng Qin, Kahou Tam, Li Li, Zijian Wang, Yuanzhe Zhao, Minglei Zhang, Chengzhong Xu
摘要
Deploying large language models (LLMs) on edge devices is crucial for delivering fast responses and ensuring data privacy. However, the limited storage, weight, and power of edge devices make it difficult to deploy LLM-powered applications. These devices must balance latency requirements with energy consumption and model accuracy. In this paper, we first quantify the challenges of deploying LLMs on off-the-shelf edge devices and then we present CLONE, an in-depth algorithm-hardware co-design at both the model- and system-level that intelligently integrates real-time, energy optimization while maintaining robust generality. In order to maximize the synergistic benefits of these algorithms in always-on and intermediate edge computing settings, we specialize in a 28nm scalable hardware accelerator system. We implement and extensively evaluate CLONE on two off-the-shelf edge platforms. Experiments show that CLONE effectively accelerates the inference process up to 11.92x, and saves energy up to 7.36x, while maintaining high-generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM InferenceJiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LINICML 2026 · 被引用 3 次
- Streaming Video Crime Anticipation with Spatio-Temporal Causal ReasoningYusong Wang, Zheyuan Gu, Keyu Mao, Minghao Shao 等CVPR 2026 · 被引用 1 次
- FARLock: Asymmetric RDMA Locking Made FairYuehao Hu, Jiatang Zhou, Tianzheng Wang, Keval VoraOSDI 2026
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 被引用 2 次
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 被引用 10 次
- CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge DeploymentQinfeng Li, Tianyue Luo, Xuhong Zhang, Yangfan Xie 等NeurIPS 2025 · 被引用 9 次
- VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible AcceleratorZhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang 等DAC 2025 · 被引用 1 次
- A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge SystemsYuze Liu, Yunhan Wang, Tiehua Zhang, Zhishu Shen 等WWW 2026 · 被引用 1 次
