Multi-dimensional Assessment of Crowdsourced Testing Reports via LLMs
Yue Wang, Yuan Zhang, Shengcheng Yu, Zhenyu Chen
Abstract
Crowdsourced testing can markedly enhance test coverage and the discovery rate of potential defects compared to traditional software testing, making it increasingly popular. However, with the widespread use of crowdsourced testing, more and more crowdworkers from various backgrounds are submitting a large number of testing reports to crowdsourced testing platforms, which hinders developers from effectively reviewing the reports. Facing a vast amount of reports with varying quality, manual review is not only time-consuming and labor-intensive but also increases costs. Therefore, how to efficiently review crowdsourced testing reports has become a major challenge. To address this challenge, we propose a multi-dimensional assessment method for crowdsourced testing reports based on large language models. This method not only inherits the textuality dimension widely used in traditional report assessment but also innovatively introduces two new dimensions: adequacy and competitiveness. It comprehensively assesses the quality of crowdsourced testing reports from multiple perspectives, aiming to better screen for high-quality crowdsourced testing reports. Through experimental analysis conducted on three different applications, we have proven the consistency of our method with human raters across various dimensions, and we have also observed an enhancement in the efficiency of report assessment.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get adcafa73-56d2-4de6-bfa4-9ec1bb20ae67Related papers
- Towards Automated Crowdsourced Testing via Personified-LLMShengcheng Yu, Yuchen Ling, Chunrong Fang, Zhenyu Chen et al.FSE 2026 · 1 citation
- Prioritize Crowdsourced Test Reports via Deep Screenshot UnderstandingShengcheng Yu, Chunrong Fang, Zhenfei Cao, Xu Wang et al.ICSE 2021 · 30 citations
- Semi-supervised Crowdsourced Test Report Clustering via Screenshot-Text Binding RulesShengcheng Yu, Chunrong Fang, Quanjun Zhang, Mingzhe Du et al.FSE 2024 · 4 citations
- Human-LLM Hybrid Text Answer Aggregation for Crowd AnnotationsJiyi LiEMNLP 2024 · 1 citation
- QG-SMS: Enhancing Test Item Analysis via Student Modeling and SimulationBang Nguyen, Tingting Du, Mengxia Yu, Lawrence Angrave et al.ACL 2025 · 2 citations
