MetaTPTrans: A Meta Learning Approach for Multilingual Code Representation Learning
Weiguo Pian, Hanyu Peng, Xunzhu Tang, Tiezhu Sun, Haoye Tian, Andrew Habib, Jacques Klein, Tegawendé F. Bissyandé
摘要
Representation learning of source code is essential for applying machine learning to software engineering tasks. Learning code representation from a multilingual source code dataset has been shown to be more effective than learning from single-language datasets separately, since more training data from multilingual dataset improves the model's ability to extract language-agnostic information from source code. However, existing multilingual training overlooks the language-specific information which is crucial for modeling source code across different programming languages, while only focusing on learning a unified model with shared parameters among different languages for language-agnostic information modeling. To address this problem, we propose MetaTPTrans, a meta learning approach for multilingual code representation learning. MetaTPTrans generates different parameters for the feature extractor according to the specific programming language type of the input code snippet, enabling the model to learn both language-agnostic and language-specific information with dynamic parameters in the feature extractor. We conduct experiments on the code summarization and code completion tasks to verify the effectiveness of our approach. The results demonstrate the superiority of our approach with significant improvements on state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CREF: An LLM-Based Conversational Software Repair Framework for Programming TutorsBoyang Yang, Haoye Tian, Weiguo Pian, Haoran Yu 等ISSTA 2024 · 被引用 26 次
- Zero-Shot Cross-Domain Code Search without Fine-TuningKeyu Liang, Zhongxin Liu, Chao Liu, Zhiyuan Wan 等FSE 2025 · 被引用 2 次
- Beyond Language Boundaries: Uncovering Programming Language Families for Code Language ModelsShangbo Yun, Xiaodong Gu, Jianghong Huang, Beijun ShenFSE 2026
- IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code GeneratorsIndraneil Paul, Goran Glavas, Iryna GurevychACL 2024
它引用的顶会 Paper7
- Directional Message Passing for Molecular GraphsJohannes Klicpera, Janek Groß, Stephan GünnemannICLR 2020 · 被引用 1,079 次
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 被引用 162 次
相关 Paper
- Language-Agnostic Representation Learning of Source Code from Structure and ContextDaniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec 等ICLR 2021 · 被引用 131 次
- Low-Resources Project-Specific Code SummarizationRui Xie, Tianxiang Hu, Wei Ye, Shikun ZhangASE 2022 · 被引用 15 次
- Integrating Tree Path in Transformer for Code RepresentationHan Peng, Ge Li, Wenhan Wang, Yunfei Zhao 等NeurIPS 2021 · 被引用 56 次
- Towards Low-Resource Automatic Program Repair with Meta-Learning and Pretrained Language ModelsWeishi Wang, Yue Wang, Steven C. H. Hoi, Shafiq JotyEMNLP 2023 · 被引用 2 次
- Meta Distant Transfer Learning for Pre-trained Language ModelsChengyu Wang, Haojie Pan, Minghui Qiu, Jun Huang 等EMNLP 2021 · 被引用 3 次
