Glider: A Reinforcement Learning Approach to Extract UI Scripts from Websites
Yuanchun Li, Oriana Riva
Abstract
Web automation scripts (tasklets) are used by personal AI assistants to carry out human tasks such as reserving a car or buying movie tickets. Generating tasklets today is a tedious job which requires much manual effort. We propose Glider, an automated and scalable approach to generate tasklets from a natural language task query and a website URL. A major advantage of Glider is that it does not require any pre-training. Glider models tasklet extraction as a state space search, where agents can explore a website's UI and get rewarded when making progress towards task completion. The reward is computed based on the agent's navigating pattern and the similarity between its trajectory and the task query. A hierarchical reinforcement learning policy is used to efficiently find the action sequences that maximize the reward. To evaluate Glider, we used it to extract tasklets for tasks in various categories (shopping, real-estate, flights, etc.); in 79% of cases a correct tasklet was generated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d3f7fcf-0f58-44d9-b87d-24c87f45d27bCited by top-tier papers6
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao et al.MobiCom 2024 · 94 citations
- MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task AutomationSunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi et al.MobiCom 2024 · 21 citations
- META-GUI: Towards Multi-modal Conversational Agents on Mobile GUILiangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai et al.EMNLP 2022 · 11 citations
- Etna: Harvesting Action Graphs from WebsitesOriana Riva, Jason KaceUIST 2021 · 6 citations
- Feature-Driven End-to-End Test GenerationParsa Alian, Noor Nashid, Mobina Shahbandeh, Taha Shabani et al.ICSE 2025 · 2 citations
Builds on2
- Translating video recordings of mobile app usages into replayable scenariosCarlos Bernal-Cárdenas, Nathan Cooper, Kevin Moran, Oscar Chaparro et al.ICSE 2020 · 61 citations
- Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in ContextShashank Srivastava, Oleksandr Polozov, Nebojsa Jojic, Christopher MeekACL 2020 · 2 citations
Related papers
- InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingZiyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo et al.ACL 2026 · 11 citations
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsYiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang et al.ICLR 2025
- Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningZican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue et al.ICML 2025
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag et al.ACL 2025 · 15 citations
- Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet AgentsYifei Zhou, Qianlan Yang, Kaixiang Lin, Min Bai et al.ICML 2025
