From Natural Language to Executable Properties for Property-Based Testing of Mobile Apps (Experience Paper)
Yiheng Xiong, Ting Su, Jingling Sun, Jue Wang, Qin Li, Geguang Pu, Zhendong Su
摘要
Property-based testing (PBT) is a popular software testing methodology and is effective in validating the functionality of mobile applications (apps for short). However, its adoption in practice remains limited, largely due to the manual effort and technical expertise required to specify executable properties. In this experience paper, we propose a novel structured property synthesis approach that automatically translates property descriptions in natural language into executable properties, and implement it in a tool named iPBT. Our approach decomposes the problem into UI semantic grounding and executable property synthesis. It first builds an enriched widget context via multimodal LLMs to align visual elements with their functional semantics, and then uses an LLM with in-context learning to generate framework-specific executable properties. We evaluate with a closed-source LLM (GPT-4o) and an open-source LLM (DeepSeek-V3) on 160 diverse property descriptions across 20 apps (124 from an existing benchmark and 36 newly authored). iPBT achieves 95.0% (152/160) accuracy on both LLMs. Notably, an ablation study reveals that the enriched widget context contributes to an absolute improvement of up to 18.1% (from 76.9% to 95.0%). A user study with 10 participants demonstrates that iPBT reduces the time required to write executable properties by 56%, suggesting substantially lower manual effort. Furthermore, evaluations on 1,520 linguistically diverse paraphrases of the original property descriptions further confirm iPBT’s robustness, achieving 88.2% accuracy on GPT-4o and 87.8% on DeepSeek-V3.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang 等ISSTA 2023 · 被引用 253 次
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang 等NeurIPS 2024 · 被引用 245 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
相关 Paper
- STARS: Static Analysis-Guided Assertion Synthesis using Large Language ModelsJialun Cao, Haoyu Wang, Haoran Yan, Ming Wen 等ISSTA 2026
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
- General and Practical Property-based Testing for Android AppsYiheng Xiong, Ting Su, Jue Wang, Jingling Sun 等ASE 2024 · 被引用 5 次
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等ICSE 2024 · 被引用 81 次
- Guardian: A Runtime Framework for LLM-Based UI ExplorationDezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu 等ISSTA 2024 · 被引用 13 次
