SeeAction: Towards Reverse Engineering How-What-Where of HCI Actions from Screencasts for UI Automation
Dehai Zhao, Zhenchang Xing, Qinghua Lu, Xiwei Xu, Liming Zhu
Abstract
UI automation is an useful technique for UI testing, bug reproduction and robotic process automation. Recording the user actions with an application assists rapid development of UI automation scripts, but existing recording techniques are intrusive, rely on OS or GUI framework accessibility support or assume specific app implementations. Reverse engineering user actions from screencasts is non-intrusive, but a key reverse-engineering step is currently missing - recognize human-understandable structured user actions ([command] [widget] [location]) from action screencasts. To fill the gap, we propose a deep learning based computer vision model which can recognize 11 commands and 11 widgets, and generate location phrases from action screencasts, through joint learning and multi-task learning. We label a large dataset with 7260 video-action pairs, which record the user interactions with Word, Zoom, Firefox, Photoshop and Windows 10 Settings. Through extensive experiments, we confirm the effectiveness and generality of our model, and demonstrate the usefulness of a screencast-to-action-script tool built upon our model for bug reproduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e2b64e1-ce78-4107-b083-d3bc9f5dd86eBuilds on13
- Unblind your apps: predicting natural-language labels for mobile GUI components by deep learningJieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu et al.ICSE 2020 · 101 citations
- Owl Eyes: Spotting UI Display Issues via Visual UnderstandingZhe Liu, Chunyang Chen, Junjie Wang, Yuekai Huang et al.ASE 2020 · 79 citations
- Translating video recordings of mobile app usages into replayable scenariosCarlos Bernal-Cárdenas, Nathan Cooper, Kevin Moran, Oscar Chaparro et al.ICSE 2020 · 61 citations
- Seenomaly: vision-based linting of GUI animation effects against design-don't guidelinesDehai Zhao, Zhenchang Xing, Chunyang Chen, Xiwei Xu et al.ICSE 2020 · 55 citations
- Data-driven accessibility repair revisited: on the effectiveness of generating labels for icons in Android appsForough Mehralian, Navid Salehnamadi, Sam MalekFSE 2021 · 47 citations
Related papers
- Read It, Don't Watch It: Captioning Bug Recordings AutomaticallySidong Feng, Mulong Xie, Yinxing Xue, Chunyang ChenICSE 2023 · 12 citations
- SeeHow: Workflow Extraction from Programming Screencasts through Action-Aware Video AnalyticsDehai Zhao, Zhenchang Xing, Xin Xia, Deheng Ye et al.ICSE 2023 · 6 citations
- Screencast Tutorial Video UnderstandingKunpeng Li, Chen Fang, Zhaowen Wang, Seokhwan Kim et al.CVPR 2020
- RoScript: a visual script driven truly non-intrusive robotic testing system for touch screen applicationsJu Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang et al.ICSE 2020 · 32 citations
- VideoAgentTrek: Computer-Use Pretraining from Unlabeled VideosDunjie Lu, Yiheng Xu, Junli Wang, Haoyuan Wu et al.ICLR 2026 · 20 citations
