What Could Possibly Go Wrong When Interacting with Proactive Smart Speakers? A Case Study Using an ESM Application
Jing Wei, Benjamin Tag, Johanne R. Trippas, Tilman Dingler, Vassilis Kostakos
Abstract
Voice user interfaces (VUIs) have made their way into people's daily lives, from voice assistants to smart speakers. Although VUIs typically just react to direct user commands, increasingly, they incorporate elements of proactive behaviors. In particular, proactive smart speakers have the potential for many applications, ranging from healthcare to entertainment; however, their usability in everyday life is subject to interaction errors. To systematically investigate the nature of errors, we designed a voice-based Experience Sampling Method (ESM) application to run on proactive speakers. We captured 1,213 user interactions in a 3-week field deployment in 13 participants' homes. Through auxiliary audio recordings and logs, we identify substantial interaction errors and strategies that users apply to overcome those errors. We further analyze the interaction timings and provide insights into the time cost of errors. We find that, even for answering simple ESMs, interaction errors occur frequently and can hamper the usability of proactive speakers and user experience. Our work also identifies multiple facets of VUIs that can be improved in terms of the timing of speech.
• Human-centered computing → Human computer interaction (HCI); Empirical studies in interaction design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported DataJing Wei, Sungdong Kim, Hyunhoon Jung, Young-Ho KimCSCW 2024 · 82 citations
- "This machine is for the aides": Tailoring Voice Assistant Design to Home Health Care WorkVince Bartle, Liam Albright, Nicola DellCHI 2023 · 19 citations
- Feasibility and Utility of Multimodal Micro Ecological Momentary Assessment on a SmartwatchHa Le, Veronika Potter, Rithika Lakshminarayanan, Varun Mishra et al.CHI 2025 · 11 citations
- Interface Matters: Exploring Human Trust in Health Information from Large Language Models via Text, Speech, and EmbodimentXin Sun, Yunjie Liu, Jos A. Bosch, Zhuying LiCSCW 2025 · 9 citations
- Understanding How to Administer Voice Surveys through Smart SpeakersJing Wei, Weiwei Jiang, Chaofan Wang, Difeng Yu et al.CSCW 2022 · 9 citations
Builds on11
- Skill Squatting Attacks on Amazon AlexaDeepak Kumar, Riccardo Paccagnella, Paul Murley, Eric Hennenfent et al.USENIX Security 2018 · 177 citations
- Dangerous Skills: Understanding and Mitigating Security Risks of Voice-Controlled Third-Party Functions on Virtual Personal Assistant SystemsNan Zhang, Xianghang Mi, Xuan Feng, XiaoFeng Wang et al.S&P 2019 · 160 citations
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- Hello There! Is Now a Good Time to Talk?: Opportune Moments for Proactive Interactions with Smart SpeakersNarae Cha, Auk Kim, Cheul Young Park, Soowon Kang et al.UbiComp 2020 · 74 citations
- The Role of Conversational Grounding in Supporting Symbiosis Between People and Digital AssistantsJanghee Cho, Emilee J. RaderCSCW 2020 · 50 citations
Related papers
- Understanding User Perceptions of Proactive Smart SpeakersJing Wei, Tilman Dingler, Vassilis KostakosUbiComp 2022 · 33 citations
- Exploring Context-Aware Mental Health Self-Tracking Using Multimodal Smart Speakers in Home EnvironmentsJieun Lim, Youngji Koh, Auk Kim, Uichin LeeCHI 2024 · 29 citations
- Interrupting for Microlearning: Understanding Perceptions and Interruptibility of Proactive Conversational Microlearning ServicesMinyeong Kim, Jiwook Lee, Youngji Koh, Chanhee Lee et al.CHI 2024 · 3 citations
- Better to Ask Than Assume: Proactive Voice Assistants' Communication Strategies That Respect User Agency in a Smart Home EnvironmentJeesun Oh, Wooseok Kim, Sungbae Kim, Hyeonjeong Im et al.CHI 2024 · 22 citations
- Beyond Words: Measuring User Experience through Speech Analysis in Voice User InterfacesYong Ma, Xuesong Zhang, Xuedong Zhang, Natalia Bartlomiejczyk et al.CHI 2026 · 1 citation
