From Natural Language to Executable Properties for Property-Based Testing of Mobile Apps (Experience Paper)
Yiheng Xiong, Ting Su, Jingling Sun, Jue Wang, Qin Li, Geguang Pu, Zhendong Su
Abstract
Property-based testing (PBT) is a popular software testing methodology and is effective in validating the functionality of mobile applications (apps for short). However, its adoption in practice remains limited, largely due to the manual effort and technical expertise required to specify executable properties. In this experience paper, we propose a novel structured property synthesis approach that automatically translates property descriptions in natural language into executable properties, and implement it in a tool named iPBT. Our approach decomposes the problem into UI semantic grounding and executable property synthesis. It first builds an enriched widget context via multimodal LLMs to align visual elements with their functional semantics, and then uses an LLM with in-context learning to generate framework-specific executable properties. We evaluate with a closed-source LLM (GPT-4o) and an open-source LLM (DeepSeek-V3) on 160 diverse property descriptions across 20 apps (124 from an existing benchmark and 36 newly authored). iPBT achieves 95.0% (152/160) accuracy on both LLMs. Notably, an ablation study reveals that the enriched widget context contributes to an absolute improvement of up to 18.1% (from 76.9% to 95.0%). A user study with 10 participants demonstrates that iPBT reduces the time required to write executable properties by 56%, suggesting substantially lower manual effort. Furthermore, evaluations on 1,520 linguistically diverse paraphrases of the original property descriptions further confirm iPBT’s robustness, achieving 88.2% accuracy on GPT-4o and 87.8% on DeepSeek-V3.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b30a444d-f3c7-4f84-a132-8f90cbbdd693Builds on23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
Related papers
- STARS: Static Analysis-Guided Assertion Synthesis using Large Language ModelsJialun Cao, Haoyu Wang, Haoran Yan, Ming Wen et al.ISSTA 2026
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che et al.ICSE 2023 · 107 citations
- General and Practical Property-based Testing for Android AppsYiheng Xiong, Ting Su, Jue Wang, Jingling Sun et al.ASE 2024 · 5 citations
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.ICSE 2024 · 81 citations
- Guardian: A Runtime Framework for LLM-Based UI ExplorationDezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu et al.ISSTA 2024 · 13 citations
