Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT Models
Jinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang, Jane Cleland-Huang
摘要
Software traceability establishes and leverages associations between diverse development artifacts. Researchers have proposed the use of deep learning trace models to link natural language artifacts, such as requirements and issue descriptions, to source code; however, their effectiveness has been restricted by availability of labeled data and efficiency at runtime. In this study, we propose a novel framework called Trace BERT (T-BERT) to generate trace links between source code and natural language artifacts. To address data sparsity, we leverage a three-step training strategy to enable trace models to transfer knowledge from a closely related Software Engineering challenge, which has a rich dataset, to produce trace links with much higher accuracy than has previously been achieved. We then apply the T-BERT framework to recover links between issues and commits in Open Source Projects. We comparatively evaluated accuracy and efficiency of three BERT architectures. Results show that a Single-BERT architecture generated the most accurate links, while a Siamese-BERT architecture produced comparable results with significantly less execution time. Furthermore, by learning and transferring knowledge, all three models in the framework outperform classical IR trace models. On the three evaluated real-word OSS projects, the best T-BERT stably outperformed the VSM model with average improvements of 60.31% measured using Mean Average Precision (MAP). RNN severely underperformed on these projects due to insufficient training data, while T-BERT overcame this problem by using pretrained language models and transfer learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao 等ICCV 2023 · 被引用 108 次
- PRCBERT: Prompt Learning for Requirement Classification using BERT-based Pretrained Language ModelsXianchang Luo, Yinxing Xue, Zhenchang Xing, Jiamou SunASE 2022 · 被引用 72 次
- Fast Changeset-based Bug Localization with BERTAgnieszka Ciborowska, Kostadin DamevskiICSE 2022 · 被引用 53 次
- Learning to Reduce False Positives in Analytic Bug DetectorsAnant Kharkar, Roshanak Zilouchian Moghaddam, Matthew Jin, Xiaoyu Liu 等ICSE 2022 · 被引用 33 次
- CCRep: Learning Code Change Representations via Pre-Trained Code Model and Query BackZhongxin Liu, Zhijie Tang, Xin Xia, Xiaohu YangICSE 2023 · 被引用 25 次
它引用的顶会 Paper2
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
相关 Paper
- EALink: An Efficient and Accurate Pre-Trained Framework for Issue-Commit Link RecoveryChenyuan Zhang, Yanlin Wang, Zhao Wei, Yong Xu 等ASE 2023 · 被引用 10 次
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationDominik Fuchß, Tobias Hey, Jan Keim, Haoyu Liu 等ICSE 2025 · 被引用 8 次
- A Comparative Study of Transformer-Based Neural Text Representation Techniques on Bug TriagingAtish Kumar Dipongkor, Kevin MoranASE 2023 · 被引用 9 次
- Semi-supervised pre-processing for learning-based traceability framework on real-world software projectsLiming Dong, He Zhang, Wei Liu, Zhiluo Weng 等FSE 2022 · 被引用 15 次
- Improving Fault Localization and Program Repair with Deep Semantic Features and Transferred KnowledgeXiangxin Meng, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2022 · 被引用 77 次
