Learning to Detect and Localize Multilingual Bugs
Haoran Yang, Yu Nong, Tao Zhang, Xiapu Luo, Haipeng Cai
Abstract
Increasing studies have shown bugs in multi-language software as a critical loophole in modern software quality assurance, especially those induced by language interactions (i.e., multilingual bugs). Yet existing tool support for bug detection/localization remains largely limited to single-language software, despite the long-standing prevalence of multi-language systems in various real-world software domains. Extant static/dynamic analysis and deep learning (DL) based approaches all face major challenges in addressing multilingual bugs. In this paper, we present xLoc, a DL-based technique/tool for detecting and localizing multilingual bugs. Motivated by results of our bug-characteristics study on top locations of multilingual bugs, xLoc first learns the general knowledge relevant to differentiating various multilingual control-flow structures. This is achieved by pre-training a Transformer model with customized position encoding against novel objectives. Then, xLoc learns task-specific knowledge for the task of multilingual bug detection/localization, through another new position encoding scheme (based on cross-language API vicinity) that allows for the model to attend particularly to control-flow constructs that bear most multilingual bugs during fine-tuning. We have implemented xLoc for Python-C software and curated a dataset of 3,770 buggy and 15,884 non-buggy Python-C samples, which enabled our extensive evaluation of xLoc against two state-of-the-art baselines: fine-tuned CodeT5 and zero-shot ChatGPT. Our results show that xLoc achieved 94.98% F1 and 87.24%@Top-1 accuracy, which are significantly (up to 162.88% and 511.75%) higher than the baselines. Ablation studies further confirmed significant contributions of each of the novel design elements in xLoc. With respective bug-location characteristics and labeled bug datasets for fine-tuning, our design may be applied to other language combinations beyond Python-C.
• Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3de823f-2178-4a2d-bbbd-1264dc77825fCited by top-tier papers4
- Finding Compiler Bugs through Cross-Language Code Generator and Differential TestingQiong Feng, Xiaotian Ma, Ziyuan Feng, Marat Akhin et al.OOPSLA 2025 · 2 citations
- Dissecting Real-World Cross-Language BugsHaoran Yang, Haipeng CaiFSE 2025 · 2 citations
- CrossPL: Systematic Evaluation of Large Language Models for Cross Programming Language Interoperating Code Generationzhanhang xiong, Dongxia Wang, Yuekang Li, Xinyuan An et al.ICLR 2026
- Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language ModelsGuangbei Yi, Yu Nong, Minzhang Li, Haipeng CaiICSE 2026
Builds on25
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- TreeGen: A Tree-Based Transformer Architecture for Code GenerationZeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun et al.AAAI 2020 · 196 citations
- Boosting coverage-based fault localization via graph-based representation learningYiling Lou, Qihao Zhu, Jinhao Dong, Xia Li et al.FSE 2021 · 157 citations
- Fault Localization with Code Coverage Representation LearningYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 120 citations
Related papers
- Towards Explorative IRBL: Combining Semantic Retrieval with LLM-Driven Iterative Code ExplorationMoumita Asad, Rafed Muhammad Yasir, Sam MalekISSTA 2026
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent AgentMehil Shah, Mohammad Masudur Rahman, Foutse KhomhICSE 2026
- Toward the Automated Localization of Buggy Mobile App UIs from Bug DescriptionsAntu Saha, Yang Song, Junayed Mahmud, Ying Zhou et al.ISSTA 2024 · 7 citations
- Pre-training Code Representation with Semantic Flow Graph for Effective Bug LocalizationYali Du, Zhongxing YuFSE 2023 · 18 citations
- Fault localization to detect co-change fixing locationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2022 · 25 citations
