Automatically Generating and Improving Voice Command Interface from Operation Sequences on Smartphones
Lihang Pan, Chun Yu, Jiahui Li, Tian Huang, Xiaojun Bi, Yuanchun Shi
摘要
Using voice commands to automate smartphone tasks (e.g., making a video call) can effectively augment the interactivity of numerous mobile apps. However, creating voice command interfaces requires a tremendous amount of effort in labeling and compiling the graphical user interface (GUI) and the utterance data. In this paper, we propose AutoVCI, a novel approach to automatically generate voice command interface (VCI) from smartphone operation sequences. The generated voice command interface has two distinct features. First, it automatically maps a voice command to GUI operations and fills in parameters accordingly, leveraging the GUI data instead of corpus or hand-written rules. Second, it launches a complementary Q&A dialogue to confirm the intention in case of ambiguity. In addition, the generated voice command interface can learn and evolve from user interactions. It accumulates the history command understanding results to annotate the user's input and improve its semantic understanding ability. We implemented this approach on Android devices and conducted a two-phase user study with 16 and 67 participants in each phase. Experimental results of the study demonstrated the practical feasibility of AutoVCI.
• Human-centered computing → Interactive systems and tools; Ubiquitous and mobile computing systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningYunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma 等UIST 2024 · 被引用 24 次
- GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and LearningMinh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li 等UIST 2024 · 被引用 17 次
- From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task AutomationYiwen Yin, Yu Mei, Chun Yu, Toby Jia-Jun Li 等CHI 2025 · 被引用 8 次
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying TogglesZongru Wu, Rui Mao, Zhiyuan Tian, Pengzhou Cheng 等CVPR 2026 · 被引用 3 次
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining ApproachSeoyoung Lee, Seobin Yoon, Seongbeen Lee, Hyesoo Kim 等UIST 2025 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Voicify Your UI: Towards Android App Control with Voice CommandsMinh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari 等UbiComp 2023 · 被引用 13 次
- Video2Action: Reducing Human Interactions in Action Annotation of App Tutorial VideosSidong Feng, Chunyang Chen, Zhenchang XingUIST 2023 · 被引用 12 次
- UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action GenerationYuanzhang Lin, Zhe Zhang, He Rui, Qingao Dong 等EMNLP 2025
- Automatic Macro Mining from Interaction Traces at ScaleForrest Huang, Gang Li, Tao Li, Yang LiCHI 2024 · 被引用 11 次
- App-Based Task Shortcuts for Virtual AssistantsDeniz Arsan, Ali Zaidi, Aravind Sagar, Ranjitha KumarUIST 2021 · 被引用 10 次
