Voicify Your UI: Towards Android App Control with Voice Commands
Minh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari, Zhenchang Xing, Chunyang Chen
摘要
Nowadays, voice assistants help users complete tasks on the smartphone with voice commands, replacing traditional touchscreen interactions when such interactions are inhibited. However, the usability of those tools remains moderate due to the problems in understanding rich language variations in human commands, along with efficiency and comprehensibility issues. Therefore, we introduce Voicify, an Android virtual assistant that allows users to interact with on-screen elements in mobile apps through voice commands. Using a novel deep learning command parser, Voicify interprets human verbal input and performs matching with UI elements. In addition, the tool can directly open a specific feature from installed applications by fetching application code information to explore the set of in-app components. Our command parser achieved 90% accuracy on the human command dataset. Furthermore, the direct feature invocation module achieves better feature coverage in comparison to Google Assistant. The user study demonstrates the usefulness of Voicify in real-world scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningYunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma 等UIST 2024 · 被引用 24 次
- GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and LearningMinh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li 等UIST 2024 · 被引用 17 次
- HandProxy: Expanding the Affordances of Speech Interfaces in Immersive Environments with a Virtual Proxy HandChen Liang, Yuxuan Liu, Martez E. Mott, Anhong GuoUbiComp 2025 · 被引用 3 次
它引用的顶会 Paper8
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White 等CHI 2021 · 被引用 145 次
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang 等ACL 2020 · 被引用 75 次
- DoThisHere: Multimodal Interaction to Improve Cross-Application Tasks on Mobile DevicesJackie (Junrui) Yang, Monica S. Lam, James A. LandayUIST 2020 · 被引用 57 次
- Latte: Use-Case and Assistive-Service Driven Automated Accessibility Testing Framework for AndroidNavid Salehnamadi, Abdulaziz Alshayban, Jun-Wei Lin, Iftekhar Ahmed 等CHI 2021 · 被引用 52 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
相关 Paper
- Automatically Generating and Improving Voice Command Interface from Operation Sequences on SmartphonesLihang Pan, Chun Yu, Jiahui Li, Tian Huang 等CHI 2022 · 被引用 17 次
- Keep it Short: A Comparison of Voice Assistants' Response BehaviorGabriel Haas, Michael Rietzler, Matt Jones, Enrico RukzioCHI 2022 · 被引用 37 次
- Spying through Your Voice Assistants: Realistic Voice Command FingerprintingDilawer Ahmed, Aafaq Sabir, Anupam DasUSENIX Security 2023
- Voicemoji: Emoji Entry Using Voice for Visually Impaired PeopleMingrui Ray Zhang, Ruolin Wang, Xuhai Xu, Qisheng Li 等CHI 2021 · 被引用 25 次
- VITAS : Guided Model-based VUI Testing of VPA AppsSuwan Li, Lei Bu, Guangdong Bai, Zhixiu Guo 等ASE 2022 · 被引用 8 次
