GUIPilot: A Consistency-Based Mobile GUI Testing Approach for Detecting Application-Specific Bugs
Ruofan Liu, Xiwen Teoh, Yun Lin, Guanjie Chen, Ruofei Ren, Denys Poshyvanyk, Jin Song Dong
Abstract
GUI testing is crucial for ensuring the reliability of mobile applications. State-of-the-art GUI testing approaches are successful in exploring more application scenarios and discovering general bugs such as application crashes. However, industrial GUI testing also needs to investigate application-specific bugs such as deviations in screen layout, widget position, or GUI transition from the GUI design mock-ups created by the application designers. These mock-ups specify the expected screens, widgets, and their respective behaviors. Validating the consistency between the GUI design and the implementation is labor-intensive and time-consuming, yet, this validation step plays an important role in industrial GUI testing. In this work, we propose , an approach for detecting inconsistencies between the mobile design and their implementations. The mobile design usually consists of design mock-ups that specify (1) the expected screen appearances (e.g., widget layouts, colors, and shapes) and (2) the expected screen behaviors, regarding how one screen can transition into another (e.g., labeled widgets with textual description). Given a design mock-up and the implementation of its application, reports both their screen inconsistencies as well as process inconsistencies. On the one hand, detects the screen inconsistencies by abstracting every screen into a widget container where each widget is represented by its position, width, height, and type. By defining the partial order of widgets and the costs of replacing, inserting, and deleting widgets in a screen, we convert the screen-matching problem into an optimizable widget alignment problem. On the other hand, we translate the specified GUI transition into stepwise actions on the mobile screen (e.g., click, long-press, input text on some widgets). To this end, we propose a visual prompt for the vision-language model to infer widget-specific actions on the screen. By this means, we can validate the presence or absence of expected transitions in the implementation. Our extensive experiments on 80 mobile applications and 160 design mock-ups show that (1) can achieve 99.8% precision and 98.6% recall in detecting screen inconsistencies, outperforming the state-of-the-art approach, such as GVT, by 66.2% and 56.6% respectively, and (2) reports zero errors in detecting process inconsistencies. Furthermore, our industrial case study on applying on a trading mobile application shows that has detected nine application bugs, and all the bugs were confirmed by the original application experts. Our code is available at https://github.com/code-philia/GUIPilot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69890d17-af16-4260-9ef4-225d479c087dCited by top-tier papers5
- WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI ElementsXiwen Teoh, Yun Lin, Duc-Minh Nguyen, Ruofei Ren et al.FSE 2026 · 1 citation
- Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive PrototypingJingyu Xiao, Yuxuan Wan, Yintong Huo, Zixin Wang et al.ASE 2025 · 1 citation
- Learning Project-wise Subsequent Code Edits via Interleaving Neural-based Induction and Tool-based DeductionChenyan Liu, Yun Lin, Yuhuan Huang, Jiaxin Chang et al.ASE 2025 · 1 citation
- From Natural Language to Executable Properties for Property-Based Testing of Mobile Apps (Experience Paper)Yiheng Xiong, Ting Su, Jingling Sun, Jue Wang et al.ISSTA 2026
- EditFlow: Benchmarking and Optimizing Code Edit Recommendation Systems via Reconstruction of Developer FlowsChenyan Liu, Yun Lin, Jiaxin Chang, Jiawei Liu et al.OOPSLA 2026
Builds on45
- Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing WebpagesYun Lin, Ruofan Liu, Dinil Mon Divakaran, Jun Yang Ng et al.USENIX Security 2021 · 164 citations
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 143 citations
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che et al.ICSE 2023 · 107 citations
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao et al.MobiCom 2024 · 94 citations
Related papers
- Vision-Based Widget Mapping for Test Migration Across Mobile Platforms: Are We There Yet?Ruihua Ji, Tingwei Zhu, Xiaoqing Zhu, Chunyang Chen et al.ASE 2023 · 2 citations
- Intention-Based GUI Test Migration for Mobile Apps using Large Language ModelsShaoheng Cao, Minxue Pan, Yuanhong Lan, Xuandong LiISSTA 2025
- Practical Non-Intrusive GUI Exploration Testing with Visual-based Robotic ArmsShengcheng Yu, Chunrong Fang, Mingzhe Du, Yuchen Ling et al.ICSE 2024 · 9 citations
- From Suspicious Signals to Crashes: Guiding Bug-Driven GUI Testing via Code-Inspired TracingMengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang et al.FSE 2026
- ChromaEyes: Detecting Inconsistencies of User Interface Elements between Light and Dark Modes of Web ApplicationsK C Shweta, Byungchul Tak, Tegawendé F. Bissyandé, Dongsun KimISSTA 2026
