Automatically Generating and Improving Voice Command Interface from Operation Sequences on Smartphones
Lihang Pan, Chun Yu, Jiahui Li, Tian Huang, Xiaojun Bi, Yuanchun Shi
Abstract
Using voice commands to automate smartphone tasks (e.g., making a video call) can effectively augment the interactivity of numerous mobile apps. However, creating voice command interfaces requires a tremendous amount of effort in labeling and compiling the graphical user interface (GUI) and the utterance data. In this paper, we propose AutoVCI, a novel approach to automatically generate voice command interface (VCI) from smartphone operation sequences. The generated voice command interface has two distinct features. First, it automatically maps a voice command to GUI operations and fills in parameters accordingly, leveraging the GUI data instead of corpus or hand-written rules. Second, it launches a complementary Q&A dialogue to confirm the intention in case of ambiguity. In addition, the generated voice command interface can learn and evolve from user interactions. It accumulates the history command understanding results to annotate the user's input and improve its semantic understanding ability. We implemented this approach on Android devices and conducted a two-phase user study with 16 and 67 participants in each phase. Experimental results of the study demonstrated the practical feasibility of AutoVCI.
• Human-centered computing → Interactive systems and tools; Ubiquitous and mobile computing systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f13579a-27bb-42e6-a6ef-60063dffc508Cited by top-tier papers5
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningYunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma et al.UIST 2024 · 24 citations
- GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and LearningMinh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li et al.UIST 2024 · 17 citations
- From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task AutomationYiwen Yin, Yu Mei, Chun Yu, Toby Jia-Jun Li et al.CHI 2025 · 8 citations
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying TogglesZongru Wu, Rui Mao, Zhiyuan Tian, Pengzhou Cheng et al.CVPR 2026 · 3 citations
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining ApproachSeoyoung Lee, Seobin Yoon, Seongbeen Lee, Hyesoo Kim et al.UIST 2025 · 2 citations
Builds on2
Related papers
- Voicify Your UI: Towards Android App Control with Voice CommandsMinh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari et al.UbiComp 2023 · 13 citations
- Video2Action: Reducing Human Interactions in Action Annotation of App Tutorial VideosSidong Feng, Chunyang Chen, Zhenchang XingUIST 2023 · 12 citations
- UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action GenerationYuanzhang Lin, Zhe Zhang, He Rui, Qingao Dong et al.EMNLP 2025
- Automatic Macro Mining from Interaction Traces at ScaleForrest Huang, Gang Li, Tao Li, Yang LiCHI 2024 · 11 citations
- App-Based Task Shortcuts for Virtual AssistantsDeniz Arsan, Ali Zaidi, Aravind Sagar, Ranjitha KumarUIST 2021 · 10 citations
