GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
Bin Xie, Rui Shao, Gongwei Chen, Kaiwen Zhou, Yinchuan Li, Jie Liu, Min Zhang, Liqiang Nie
摘要
GUI automation faces critical challenges in dynamic environments. MLLMs suffer from two key issues: misinterpreting UI components and outdated knowledge. Traditional fine-tuning methods are costly for app-specific knowledge updates. We propose GUI-explorer, a trainingfree GUI agent that incorporates two fundamental mechanisms: (1) Autonomous Exploration of Function-aware Trajectory. To comprehensively cover all application functionalities, we design a Function-aware Task Goal Generator that automatically constructs exploration goals by analyzing GUI structural information (e.g., screenshots and activity hierarchies). This enables systematic exploration to collect diverse trajectories. (2) Unsupervised Mining of Transition-aware Knowledge. To establish precise screen-operation logic, we develop a Transition-aware Knowledge Extractor that extracts effective screen-operation logic through unsupervised analysis the state transition of structured interaction triples (observation, action, outcome). This eliminates the need for human involvement in knowledge extraction. With a task success rate of 53.7% on SPA-Bench and 47.4% on AndroidWorld, GUI-explorer shows significant improvements over SOTA agents. It requires no parameter updates for new apps. GUI-explorer is open-sourced and publicly available at https: //github.com/JiuTian-VL/GUI-explorer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- AgentOCR: Reimagining Agent History via Optical Self-CompressionLang Feng, Fuchao Yang, Feng Chen, Xin Cheng 等ACL 2026 · 被引用 17 次
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsXurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li 等CVPR 2026 · 被引用 11 次
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong 等CVPR 2026 · 被引用 8 次
- Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple ActionsGuo Gan, Yuxuan Ding, Cong Chen, Yuwei Ren 等ACL 2026 · 被引用 6 次
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic ManipulationYijie Zhu, Rui Shao, Ziyang Liu, Jie He 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper18
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri 等NeurIPS 2024 · 被引用 239 次
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer ControlLongtao Zheng, Rundong Wang, Xinrun Wang, Bo AnICLR 2024 · 被引用 132 次
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao 等MobiCom 2024 · 被引用 94 次
相关 Paper
- M-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data MiningRui Lyu, Juncheng Mo, Tianyi Chu, Chen Rao 等ICLR 2026
- GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous ExplorationYue Fan, Handong Zhao, Ruiyi Zhang, Yu Shen 等EMNLP 2025 · 被引用 15 次
- Agent-SAMA: State-Aware Mobile AssistantLinqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun (Peter) Chen 等AAAI 2026 · 被引用 2 次
- GUI-Xplore: Empowering Generalizable GUI Agents with One ExplorationYuchen Sun, Shanhui Zhao, Tao Yu, Hao Wen 等CVPR 2025
- Scaling Synthetic Task Generation for Agents via ExplorationRam Ramrakhya, Andrew Szot, Omar Attia, Bogdan Mazoure 等ICLR 2026 · 被引用 15 次
