MetaTPTrans: A Meta Learning Approach for Multilingual Code Representation Learning
Weiguo Pian, Hanyu Peng, Xunzhu Tang, Tiezhu Sun, Haoye Tian, Andrew Habib, Jacques Klein, Tegawendé F. Bissyandé
Abstract
Representation learning of source code is essential for applying machine learning to software engineering tasks. Learning code representation from a multilingual source code dataset has been shown to be more effective than learning from single-language datasets separately, since more training data from multilingual dataset improves the model's ability to extract language-agnostic information from source code. However, existing multilingual training overlooks the language-specific information which is crucial for modeling source code across different programming languages, while only focusing on learning a unified model with shared parameters among different languages for language-agnostic information modeling. To address this problem, we propose MetaTPTrans, a meta learning approach for multilingual code representation learning. MetaTPTrans generates different parameters for the feature extractor according to the specific programming language type of the input code snippet, enabling the model to learn both language-agnostic and language-specific information with dynamic parameters in the feature extractor. We conduct experiments on the code summarization and code completion tasks to verify the effectiveness of our approach. The results demonstrate the superiority of our approach with significant improvements on state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f603826-376d-463c-9490-83699fbf5a2bCited by top-tier papers4
- CREF: An LLM-Based Conversational Software Repair Framework for Programming TutorsBoyang Yang, Haoye Tian, Weiguo Pian, Haoran Yu et al.ISSTA 2024 · 26 citations
- Zero-Shot Cross-Domain Code Search without Fine-TuningKeyu Liang, Zhongxin Liu, Chao Liu, Zhiyuan Wan et al.FSE 2025 · 2 citations
- Beyond Language Boundaries: Uncovering Programming Language Families for Code Language ModelsShangbo Yun, Xiaodong Gu, Jianghong Huang, Beijun ShenFSE 2026
- IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code GeneratorsIndraneil Paul, Goran Glavas, Iryna GurevychACL 2024
Builds on7
- Directional Message Passing for Molecular GraphsJohannes Klicpera, Janek Groß, Stephan GünnemannICLR 2020 · 1,079 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 162 citations
Related papers
- Language-Agnostic Representation Learning of Source Code from Structure and ContextDaniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec et al.ICLR 2021 · 131 citations
- Low-Resources Project-Specific Code SummarizationRui Xie, Tianxiang Hu, Wei Ye, Shikun ZhangASE 2022 · 15 citations
- Integrating Tree Path in Transformer for Code RepresentationHan Peng, Ge Li, Wenhan Wang, Yunfei Zhao et al.NeurIPS 2021 · 56 citations
- Towards Low-Resource Automatic Program Repair with Meta-Learning and Pretrained Language ModelsWeishi Wang, Yue Wang, Steven C. H. Hoi, Shafiq JotyEMNLP 2023 · 2 citations
- Meta Distant Transfer Learning for Pre-trained Language ModelsChengyu Wang, Haojie Pan, Minghui Qiu, Jun Huang et al.EMNLP 2021 · 3 citations
