Rataplan: Resilient Automation of User Interface Actions with Multi-modal Proxies
Tom Veuskens, Kris Luyten, Raf Ramakers
Abstract
We present Rataplan, a robust and resilient pixel-based approach for linking multi-modal proxies to automated sequences of actions in graphical user interfaces (GUIs). With Rataplan, users demonstrate a sequence of actions and answer human-readable follow-up questions to clarify their desire for automation. After demonstrating a sequence, the user can link a proxy input control to the action which can then be used as a shortcut for automating a sequence. Alternatively, output proxies use a notification model in which content is pushed when it becomes available. As an example use case, Rataplan uses keyboard shortcuts and tangible user interfaces (TUIs) as input proxies, and TUIs as output proxies. Instead of relying on available APIs, Rataplan automates GUIs using pixel-based reverse engineering. This ensures our approach can be used with all applications that offer a GUI, including web applications. We implemented a set of important strategies to support robust automation of modern interfaces that have a flat and minimal style, have frequent data and state changes, and have dynamic viewports.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 20bdd8d4-513c-403b-a1c7-252249550a68Related papers
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang et al.ICLR 2026 · 109 citations
- SLAP: Shortcut Learning for Abstract PlanningY. Isabel Liu, Bowen Li, Benjamin Eysenbach, Tom SilverICLR 2026 · 6 citations
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining ApproachSeoyoung Lee, Seobin Yoon, Seongbeen Lee, Hyesoo Kim et al.UIST 2025 · 2 citations
- Appliancizer: Transforming Web Pages into Electronic DevicesJorge Garza, Devon J. Merrill, Steven SwansonCHI 2021 · 7 citations
- From Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesPeter Shaw, Mandar Joshi, James Cohan, Jonathan Berant et al.NeurIPS 2023 · 89 citations
