Sorry, I don't Understand: Improving Voice User Interface Testing
Emanuela Guglielmi, Giovanni Rosa, Simone Scalabrino, Gabriele Bavota, Rocco Oliveto
摘要
Voice-based virtual assistants are becoming increasingly popular. Such systems provide frameworks to developers on which they can build their own apps. End-users can interact with such apps through a Voice User Interface (VUI), which allows to use natural language commands to perform actions. Testing such apps is far from trivial: The same command can be expressed in different ways. To support developers in testing VUIs, Deep Learning (DL)-based tools have been integrated in the development environments (e.g., the Alexa Developer Console, or ADC) to generate paraphrases for the commands (seed utterances) specified by the developers. Such tools, however, generate few paraphrases that do not always cover corner cases. In this paper, we introduce VUI-UPSET, a novel approach that aims at adapting chatbot-testing approaches to VUI-testing. Both systems, indeed, provide a similar natural-language-based interface to users. We conducted an empirical study to understand how VUI-UPSET compares to existing approaches in terms of (i) correctness of the generated paraphrases, and (ii) capability of revealing bugs. Multiple authors analyzed 5,872 generated paraphrases, with a total of 13,310 manual evaluations required for such a process. Our results show that, while the DL-based tool integrated in the ADC generates a higher percentage of meaningful paraphrases compared to VUI-UPSET, VUI-UPSET generates more bug-revealing paraphrases. This allows developers to test more thoroughly their apps at the cost of discarding a higher number of irrelevant paraphrases.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- VITAS : Guided Model-based VUI Testing of VPA AppsSuwan Li, Lei Bu, Guangdong Bai, Zhixiu Guo 等ASE 2022 · 被引用 8 次
- Life after Speech Recognition: Fuzzing Semantic Misinterpretation for Voice Assistant ApplicationsYangyong Zhang, Lei Xu, Abner Mendoza, Guangliang Yang 等NDSS 2019 · 被引用 60 次
- SkillDetective: Automated Policy-Violation Detection of Voice Assistant Applications in the WildJeffrey Young, Song Liao, Long Cheng, Hongxin Hu 等USENIX Security 2022
- Voice App Developer Experiences with Alexa and Google Assistant: Juggling Risks, Liability, and SecurityWilliam Seymour, Noura Abdi, Kopo M. Ramokapane, Jide S. Edu 等USENIX Security 2024 · 被引用 9 次
- Scrutinizing Privacy Policy Compliance of Virtual Personal Assistant AppsFuman Xie, Yanjun Zhang, Chuan Yan, Suwan Li 等ASE 2022 · 被引用 31 次
