GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
Minh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li, Shengdong Zhao, Zhenchang Xing, Chunyang Chen
Abstract
Virtual assistants have the potential to play an important role in helping users achieves different tasks. However, these systems face challenges in their real-world usability, characterized by inefficiency and struggles in grasping user intentions. Leveraging recent advances in Large Language Models (LLMs), we introduce GptVoiceTasker, a virtual assistant poised to enhance user experiences and task efficiency on mobile devices. GptVoiceTasker excels at intelligently deciphering user commands and executing relevant device interactions to streamline task completion. For unprecedented tasks, GptVoiceTasker utilises the contextual information and on-screen content to continuously explore and execute the tasks. In addition, the system continually learns from historical user commands to automate subsequent task invocations, further enhancing execution efficiency. From our experiments, GptVoiceTasker achieved 84.5% accuracy in parsing human commands into executable actions and 85.7% accuracy in automating multi-step tasks. In our user study, GptVoiceTasker boosted task efficiency in real-world scenarios by 34.85%, accompanied by positive participant feedback. We made GptVoiceTasker open-source, inviting further research into LLMs utilization for diverse tasks through prompt engineering and leveraging user usage data to improve efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d10df4e-284d-4978-a337-ae7b161aeee0Cited by top-tier papers9
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.ACM MM 2024 · 15 citations
- Deaf and Hard of Hearing Access to Intelligent Personal Assistants: Comparison of Voice-Based Options with an LLM-Powered Touch InterfacePaige S. DeVries, Michaela Okosi, Ming Li, Nora Dunphy et al.CHI 2026 · 2 citations
- Understanding Spatiotemporal-Aware Multimodal Conversational Search in the Outdoor Urban SpaceJiangnan Xu, Suyeon Seo, Joni Salminen, Michael Saker et al.CHI 2026 · 1 citation
- TaskAudit: Detecting Functiona11ity Errors in Mobile Apps via Agentic Task ExecutionMingyuan Zhong, Xia Chen, Davin Win Kyi, Chen Li et al.CHI 2026 · 1 citation
- CHAMELEOSCAN: Demystifying and Detecting iOS Chameleon Apps via LLM-Powered UI ExplorationHongyu Lin, Yicheng Hu, Haitao Xu, Yanchen Lu et al.NDSS 2026 · 1 citation
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
Related papers
- VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningYunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma et al.UIST 2024 · 24 citations
- MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task AutomationSunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi et al.MobiCom 2024 · 21 citations
- Voicify Your UI: Towards Android App Control with Voice CommandsMinh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari et al.UbiComp 2023 · 13 citations
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao et al.MobiCom 2024 · 94 citations
- Supporting Text Entry in Virtual Reality with Large Language ModelsLiuqing Chen, Yu Cai, Ruyue Wang, Shixian Ding et al.IEEE VR 2024 · 18 citations
