Semi-supervised Crowdsourced Test Report Clustering via Screenshot-Text Binding Rules
Shengcheng Yu, Chunrong Fang, Quanjun Zhang, Mingzhe Du, Jia Liu, Zhenyu Chen
Abstract
Due to the openness of the crowdsourced testing paradigm, crowdworkers submit massive spotty duplicate test reports, which hinders developers from effectively reviewing the reports and detecting bugs. Test report clustering is widely used to alleviate this problem and improve the effectiveness of crowdsourced testing. Existing clustering methods basically rely on the analysis of textual descriptions. A few methods are independently supplemented by analyzing screenshots in test reports as pixel sets, leaving out the semantics of app screenshots from the widget perspective. Further, ignoring the semantic relationships between screenshots and textual descriptions may lead to the imprecise analysis of test reports, which in turn negatively affects the clustering effectiveness. This paper proposes a semi-supervised crowdsourced test report clustering approach, namely SemCluster . SemCluster respectively extracts features from app screenshots and textual descriptions and forms the structure feature, the content feature, the bug feature, and reproduction steps. The clustering is principally conducted on the basis of the four features. Further, in order to avoid bias of specific individual features, SemCluster exploits the semantic relationships between app screenshots and textual descriptions to form the semantic binding rules as guidance for clustering crowdsourced test reports. Experiment results show that SemCluster outperforms state-of-the-art approaches on six widely used metrics by 10.49% – 200.67%, illustrating the excellent effectiveness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7e6060b8-c0d7-49f3-89bb-3c412ad152fdCited by top-tier papers1
Ask how each one uses itRelated papers
- Prioritize Crowdsourced Test Reports via Deep Screenshot UnderstandingShengcheng Yu, Chunrong Fang, Zhenfei Cao, Xu Wang et al.ICSE 2021 · 30 citations
- Multi-dimensional Assessment of Crowdsourced Testing Reports via LLMsYue Wang, Yuan Zhang, Shengcheng Yu, Zhenyu ChenASE 2025
- It Takes Two to TANGO: Combining Visual and Textual Information for Detecting Duplicate Video-Based Bug ReportsNathan Cooper, Carlos Bernal-Cárdenas, Oscar Chaparro, Kevin Moran et al.ICSE 2021 · 2 citations
- Semantic matching of GUI events for test reuse: are we there yet?Leonardo Mariani, Ali Mohebbi, Mauro Pezzè, Valerio TerragniISSTA 2021 · 34 citations
- Standing on the Shoulders of Giants: Bug-Aware Automated GUI Testing via Retrieval AugmentationMengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang et al.FSE 2025 · 4 citations
