VCU-LLM: Prompt-efficient On-device Large Language Model for Vague Command Understanding in Smart Homes
Zhengyuan Zhang, Dong Zhao, Tiancheng He, Zilong Wang, Xiangyu Li, Huadong Ma
摘要
Smart homes are receiving growing interest in global markets. However, current smart home assistant systems either control appliances solely through users' explicit commands or rely on cloud-based large language models (LLMs) for vague command understanding, which brings drawbacks such as high latency, privacy concern, and high cost. In this work, we propose VCU-LLM, the first system to deploy LLMs on edge devices for local vague command understanding and smart device control plan generation. VCU-LLM introduces a novel vague command knowledge retrieval algorithm that refines device-related information in the input prompt, thereby accelerating the LLM's on-device inference and reducing task complexity. We further construct a dataset for LLM fine-tuning to simulate the use of smart home assistants in controlling devices across different households. During inference, a customized KV-cache technique is applied for further inference acceleration. Our evaluations with both human-based and LLM-based scoring demonstrate that VCU-LLM improves the quality of generated control plans by an average of 43.3% compared with SOTA baselines, while reducing time overhead by an average of 8.44x compared with other on-device baselines. We also implement VCU-LLM through a case study in a real home environment, demonstrating its feasibility in real-world application.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 被引用 10 次
- Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language ModelsEvan King, Haoxiang Yu, Sangsu Lee, Christine JulienUbiComp 2024 · 被引用 91 次
- HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple DevicesSilin Li, Yuhang Guo, Jiashu Yao, Zeming Liu 等ACL 2025
- AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the EdgeTao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou 等INFOCOM 2025 · 被引用 5 次
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 被引用 2 次
