Semi-supervised pre-processing for learning-based traceability framework on real-world software projects
Liming Dong, He Zhang, Wei Liu, Zhiluo Weng, Hongyu Kuang
Abstract
The traceability of software artifacts has been recognized as an important factor to support various activities in software development processes. However, traceability can be difficult and time-consuming to create and maintain manually, thereby automated approaches have gained much attention. Unfortunately, existing automated approaches for traceability suffer from practical issues. This paper aims to gain an understanding of the potential challenges for the underperforming of the state-of-the-art, ML-based trace link classifiers applied in real-world projects. By investigating different industrial datasets, we found that two critical (and classic) challenges, i.e. data imbalance and sparse problems, lie in real-world projects’ traceability automation. To overcome these challenges, we developed a framework called SPLINT to incorporate hybrid textual similarity measures and semi-supervised learning strategies as enhancements to the learning-based traceability approaches. We carried out experiments with six open-source platforms and ten industry datasets. The results confirm that SPLINT is able to operate at higher performance on two communities’ datasets. Specifically, the industrial datasets, which significantly suffer from data imbalance and sparsity problems, show an increase in F2-score over 14% and AUC over 8% on average. The adjusted class-balancing and self-training policies used in SPLINT (CBST-Adjust) also work effectively for the selection of pseudo-labels on minor classes from unlabeled trace sets, demonstrating SPLINT’s practicability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c85dab84-db04-4369-949a-215918d68484Cited by top-tier papers5
- EALink: An Efficient and Accurate Pre-Trained Framework for Issue-Commit Link RecoveryChenyuan Zhang, Yanlin Wang, Zhao Wei, Yong Xu et al.ASE 2023 · 10 citations
- TRIAD: Automated Traceability Recovery based on Biterm-enhanced Deduction of Transitive Links among ArtifactsHui Gao, Hongyu Kuang, Wesley K. G. Assunção, Christoph Mayr-Dorn et al.ICSE 2024 · 8 citations
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationDominik Fuchß, Tobias Hey, Jan Keim, Haoyu Liu et al.ICSE 2025 · 8 citations
- LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link RecoveryArshia Akhavan, Alireza Hoseinpour, Abbas Heydarnoori, Hamid Bagheri et al.FSE 2026
- Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted RetrievalHuihui Huang, Ratnadira Widyasari, Ting Zhang, Ivana Clairine Irsan et al.ICSE 2026
Related papers
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang et al.ICSE 2021 · 124 citations
- Improving the effectiveness of traceability link recovery using hierarchical bayesian networksKevin Moran, David N. Palacio, Carlos Bernal-Cárdenas, Daniel McCrystal et al.ICSE 2020 · 40 citations
- Establishing multilevel test-to-code traceability linksRobert White, Jens Krinke, Raymond TanICSE 2020 · 37 citations
- Using Consensual Biterms from Text Structures of Requirements and Code to Improve IR-Based Traceability RecoveryHui Gao, Hongyu Kuang, Kexin Sun, Xiaoxing Ma et al.ASE 2022 · 20 citations
- Integrating Multiple Features for Weakly-Supervised False-Passing Products Detection in Software Product LinesTao Zhang, Yan Lei, Haoran Xia, Huan Xie et al.ISSTA 2026
