META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, Kai Yu
Abstract
Task-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation. Current TOD systems usually focus on multi-turn text/speech interaction, then they would call back-end APIs designed for TODs to perform the task. However, this API-based architecture greatly limits the information-searching capability of intelligent assistants and may even lead to task failure if TOD-specific APIs are not available or the task is too complicated to be executed by the provided APIs. In this paper, we propose a new TOD architecture: GUIbased task-oriented dialogue system (GUI-TOD). A GUI-TOD system can directly perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs. Furthermore, we release META-GUI, a dataset for training a Multi-modal convErsaTional Agent on mobile GUI. We also propose a multi-model action prediction and response model, which show promising results on META-GUI. The dataset, codes and leaderboard are publicly available ‡ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50465b4c-ae08-4584-99ee-6397c6ce138eCited by top-tier papers28
- WebLINX: Real-World Website Navigation with Multi-Turn DialogueXing Han Lù, Zdenek Kasner, Siva ReddyICML 2024 · 146 citations
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao et al.MobiCom 2024 · 94 citations
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng et al.ACL 2025 · 71 citations
- Large Language Models Empowered Personalized Web AgentsHongru Cai, Yongqi Li, Wenjie Wang, Fengbin Zhu et al.WWW 2025 · 62 citations
- Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental DistractionsXinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan et al.ACL 2025 · 54 citations
Builds on7
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- ActionBert: Leveraging User Actions for Semantic Understanding of User InterfacesZecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu et al.AAAI 2021 · 91 citations
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded DialoguesHung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. HoiEMNLP 2020 · 30 citations
- Glider: A Reinforcement Learning Approach to Extract UI Scripts from WebsitesYuanchun Li, Oriana RivaSIGIR 2021 · 11 citations
Related papers
- Rethinking Task-Oriented Dialogue Systems: From Complex Modularity to Zero-Shot Autonomous AgentHeng-Da Xu, Xian-Ling Mao, Puhai Yang, Fanshu Sun et al.ACL 2024
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu et al.EMNLP 2024
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsBofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang et al.AAAI 2026 · 26 citations
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 210 citations
