Learning without Forgetting: Towards Continual learning of Fault Localization Models in Industrial Software Systems
Chun Li, Hui Li, Zhong Li, Minxue Pan, Xuandong Li
摘要
Learning-based fault localization has achieved promising results. However, as software and tests are constantly evolving, models trained on old data become ineffective on new data. Particularly, in the context of system testing for large-scale software, each iteration generates a large volume of new data. This makes retraining the model from scratch incur an unacceptable time overhead, while merely fine-tuning on new data leads to catastrophic forgetting. Continual learning offers an effective method for models to avoid catastrophic forgetting during this iterative process. However, existing continual learning methods are not specifically designed for fault localization or for large-scale software system testing scenarios, which leads to their direct application yielding sub-optimal effectiveness. In response, we propose Ciallo, a novel continual learning framework specifically designed for large-scale software fault localization. Ciallo first extracts fine-grained program semantics from logs, then utilizes fault characteristics to enhance the weights of certain semantics. Finally, Ciallo uses an unsupervised algorithm to obtain corresponding embeddings and selects representative exemplars based on clustering. Subsequently, Ciallo mixes the representative exemplars with new samples for training and adjusts the loss weight according to the model’s degree of mastery over the sample. This allows the model to focus more on samples that are not yet well-mastered during the training process, thereby enabling it to learn new faults while mitigating the forgetting of old ones. In extensive evaluations against 6 continual learning baselines, Ciallo demonstrates superior performance, improving overall effectiveness by 17.30% to 45.23%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Enhancing Fault Localization in Industrial Software Systems via Contrastive LearningChun Li, Hui Li, Zhong Li, Minxue Pan 等ICSE 2025 · 被引用 1 次
- Improving Graph Learning-Based Fault Localization with Tailored Semi-supervised LearningChun Li, Hui Li, Zhong Li, Minxue Pan 等FSE 2025 · 被引用 2 次
- Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence ModelsShuzheng Gao, Hongyu Zhang, Cuiyun Gao, Chaozheng WangICSE 2023 · 被引用 14 次
- Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized RehearsalJianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang 等ACL 2024 · 被引用 13 次
- DFIL: Deepfake Incremental Learning by Exploiting Domain-invariant Forgery CluesKun Pan, Yifang Yin, Yao Wei, Feng Lin 等ACM MM 2023 · 被引用 35 次
