META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, Kai Yu
摘要
Task-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation. Current TOD systems usually focus on multi-turn text/speech interaction, then they would call back-end APIs designed for TODs to perform the task. However, this API-based architecture greatly limits the information-searching capability of intelligent assistants and may even lead to task failure if TOD-specific APIs are not available or the task is too complicated to be executed by the provided APIs. In this paper, we propose a new TOD architecture: GUIbased task-oriented dialogue system (GUI-TOD). A GUI-TOD system can directly perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs. Furthermore, we release META-GUI, a dataset for training a Multi-modal convErsaTional Agent on mobile GUI. We also propose a multi-model action prediction and response model, which show promising results on META-GUI. The dataset, codes and leaderboard are publicly available ‡ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- WebLINX: Real-World Website Navigation with Multi-Turn DialogueXing Han Lù, Zdenek Kasner, Siva ReddyICML 2024 · 被引用 146 次
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao 等MobiCom 2024 · 被引用 94 次
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng 等ACL 2025 · 被引用 71 次
- Large Language Models Empowered Personalized Web AgentsHongru Cai, Yongqi Li, Wenjie Wang, Fengbin Zhu 等WWW 2025 · 被引用 62 次
- Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental DistractionsXinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan 等ACL 2025 · 被引用 54 次
它引用的顶会 Paper7
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- ActionBert: Leveraging User Actions for Semantic Understanding of User InterfacesZecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu 等AAAI 2021 · 被引用 91 次
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
- BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded DialoguesHung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. HoiEMNLP 2020 · 被引用 30 次
- Glider: A Reinforcement Learning Approach to Extract UI Scripts from WebsitesYuanchun Li, Oriana RivaSIGIR 2021 · 被引用 11 次
相关 Paper
- Rethinking Task-Oriented Dialogue Systems: From Complex Modularity to Zero-Shot Autonomous AgentHeng-Da Xu, Xian-Ling Mao, Puhai Yang, Fanshu Sun 等ACL 2024
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu 等EMNLP 2024
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsBofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang 等AAAI 2026 · 被引用 26 次
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 被引用 210 次
