Multilingual training for Software Engineering
Toufique Ahmed, Premkumar T. Devanbu
Abstract
Well-trained machine-learning models, which leverage large amounts of open-source software data, have now become an interesting approach to automating many software engineering tasks. Several SE tasks have all been subject to this approach, with performance gradually improving over the past several years with better models and training methods. More, and more diverse, clean, labeled data is better for training; but constructing good-quality datasets is time-consuming and challenging. Ways of augmenting the volume and diversity of clean, labeled data generally have wide applicability. For some languages (e.g., Ruby) labeled data is less abundant; in others (e.g., JavaScript) the available data maybe more focused on some application domains, and thus less diverse. As a way around such data bottlenecks, we present evidence suggesting that human-written code in different languages (which performs the same function), is rather similar, and particularly preserving of identifier naming patterns; we further present evidence suggesting that identifiers are a very important element of training data for software engineering tasks. We leverage this rather fortuitous phenomenon to find evidence that available multilingual training data (across different languages) can be used to amplify performance. We study this for 3 different tasks: code summarization, code retrieval, and function naming. We note that this dataaugmenting approach is broadly compatible with different tasks, languages, and machine-learning models. CCS CONCEPTS • Software and its engineering → Software notations and tools; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e6124dd-cee4-4dbe-a9f9-5b717d9aed2fCited by top-tier papers16
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)Toufique Ahmed, Kunal Suresh Pai, Premkumar T. Devanbu, Earl T. BarrICSE 2024 · 71 citations
- CIRCLE: continual repair across programming languagesWei Yuan, Quanjun Zhang, Tieke He, Chunrong Fang et al.ISSTA 2022 · 57 citations
- One Adapter for All Programming Languages? Adapter Tuning for Code Search and SummarizationDeze Wang, Boxing Chen, Shanshan Li, Wei Luo et al.ICSE 2023 · 43 citations
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMsFederico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger et al.OOPSLA 2024 · 33 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
Related papers
- MetaTPTrans: A Meta Learning Approach for Multilingual Code Representation LearningWeiguo Pian, Hanyu Peng, Xunzhu Tang, Tiezhu Sun et al.AAAI 2023 · 18 citations
- Automating Code-Related Tasks Through Transformers: The Impact of Pre-trainingRosalia Tufano, Luca Pascarella, Gabriele BavotaICSE 2023 · 15 citations
- Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningWei Ye, Rui Xie, Jinglei Zhang, Tianxiang Hu et al.WWW 2020 · 83 citations
- Language-Agnostic Representation Learning of Source Code from Structure and ContextDaniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec et al.ICLR 2021 · 131 citations
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang et al.FSE 2022 · 65 citations
