Rataplan: Resilient Automation of User Interface Actions with Multi-modal Proxies
Tom Veuskens, Kris Luyten, Raf Ramakers
摘要
We present Rataplan, a robust and resilient pixel-based approach for linking multi-modal proxies to automated sequences of actions in graphical user interfaces (GUIs). With Rataplan, users demonstrate a sequence of actions and answer human-readable follow-up questions to clarify their desire for automation. After demonstrating a sequence, the user can link a proxy input control to the action which can then be used as a shortcut for automating a sequence. Alternatively, output proxies use a notification model in which content is pushed when it becomes available. As an example use case, Rataplan uses keyboard shortcuts and tangible user interfaces (TUIs) as input proxies, and TUIs as output proxies. Instead of relying on available APIs, Rataplan automates GUIs using pixel-based reverse engineering. This ensures our approach can be used with all applications that offer a GUI, including web applications. We implemented a set of important strategies to support robust automation of modern interfaces that have a flat and minimal style, have frequent data and state changes, and have dynamic viewports.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang 等ICLR 2026 · 被引用 109 次
- SLAP: Shortcut Learning for Abstract PlanningY. Isabel Liu, Bowen Li, Benjamin Eysenbach, Tom SilverICLR 2026 · 被引用 6 次
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining ApproachSeoyoung Lee, Seobin Yoon, Seongbeen Lee, Hyesoo Kim 等UIST 2025 · 被引用 2 次
- Appliancizer: Transforming Web Pages into Electronic DevicesJorge Garza, Devon J. Merrill, Steven SwansonCHI 2021 · 被引用 7 次
- From Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesPeter Shaw, Mandar Joshi, James Cohan, Jonathan Berant 等NeurIPS 2023 · 被引用 89 次
