HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
Silin Li, Yuhang Guo, Jiashu Yao, Zeming Liu, Haifeng Wang
摘要
Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home environment. While recent studies have explored integrating LLMs into smart home systems, they primarily focus on handling straightforward, valid single-device operation instructions. However, real-world scenarios are far more complex and often involve users issuing invalid instructions or controlling multiple devices simultaneously. These have two main challenges: LLMs must accurately identify and rectify errors in user instructions and execute multiple user instructions perfectly. To address these challenges and advance the development of LLM-based smart home assistants, we introduce HomeBench, the first smart home dataset with valid and invalid instructions across single and multiple devices in this paper. We have experimental results on 13 distinct LLMs; e.g., GPT-4o achieves only a 0.0% success rate in the scenario of invalid multi-device instructions, revealing that the existing stateof-the-art LLMs still cannot perform well in this situation even with the help of in-context learning, retrieval-augmented generation, and fine-tuning. Our code and dataset are publicly available at the link 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Towards Conversational Recommendation over Multi-Type DialogsZeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu 等ACL 2020 · 被引用 157 次
- Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language ModelsEvan King, Haoxiang Yu, Sangsu Lee, Christine JulienUbiComp 2024 · 被引用 91 次
- FEDLEGAL: The First Real-World Federated Learning Benchmark for Legal NLPZhuo Zhang, Xiangjing Hu, Jingyuan Zhang, Yating Zhang 等ACL 2023 · 被引用 18 次
相关 Paper
- Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?Guijin Son, Sangwon Baek, Sangdae Nam, Ilgyun Jeong 等ACL 2024
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia 等EMNLP 2024 · 被引用 2 次
- VCU-LLM: Prompt-efficient On-device Large Language Model for Vague Command Understanding in Smart HomesZhengyuan Zhang, Dong Zhao, Tiancheng He, Zilong Wang 等UbiComp 2026
- LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model ProgramsYunsheng Ma, Can Cui, Xu Cao, Wenqian Ye 等CVPR 2024 · 被引用 39 次
- Can Large Language Models Understand Real-World Complex Instructions?Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen 等AAAI 2024 · 被引用 99 次
