HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
Silin Li, Yuhang Guo, Jiashu Yao, Zeming Liu, Haifeng Wang
Abstract
Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home environment. While recent studies have explored integrating LLMs into smart home systems, they primarily focus on handling straightforward, valid single-device operation instructions. However, real-world scenarios are far more complex and often involve users issuing invalid instructions or controlling multiple devices simultaneously. These have two main challenges: LLMs must accurately identify and rectify errors in user instructions and execute multiple user instructions perfectly. To address these challenges and advance the development of LLM-based smart home assistants, we introduce HomeBench, the first smart home dataset with valid and invalid instructions across single and multiple devices in this paper. We have experimental results on 13 distinct LLMs; e.g., GPT-4o achieves only a 0.0% success rate in the scenario of invalid multi-device instructions, revealing that the existing stateof-the-art LLMs still cannot perform well in this situation even with the help of in-context learning, retrieval-augmented generation, and fine-tuning. Our code and dataset are publicly available at the link 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Towards Conversational Recommendation over Multi-Type DialogsZeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu et al.ACL 2020 · 157 citations
- Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language ModelsEvan King, Haoxiang Yu, Sangsu Lee, Christine JulienUbiComp 2024 · 91 citations
- FEDLEGAL: The First Real-World Federated Learning Benchmark for Legal NLPZhuo Zhang, Xiangjing Hu, Jingyuan Zhang, Yating Zhang et al.ACL 2023 · 18 citations
Related papers
- Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?Guijin Son, Sangwon Baek, Sangdae Nam, Ilgyun Jeong et al.ACL 2024
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia et al.EMNLP 2024 · 2 citations
- VCU-LLM: Prompt-efficient On-device Large Language Model for Vague Command Understanding in Smart HomesZhengyuan Zhang, Dong Zhao, Tiancheng He, Zilong Wang et al.UbiComp 2026
- LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model ProgramsYunsheng Ma, Can Cui, Xu Cao, Wenqian Ye et al.CVPR 2024 · 39 citations
- Can Large Language Models Understand Real-World Complex Instructions?Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen et al.AAAI 2024 · 99 citations
