Beyond Language Boundaries: Uncovering Programming Language Families for Code Language Models
Shangbo Yun, Xiaodong Gu, Jianghong Huang, Beijun Shen
摘要
The rapid proliferation of diverse programming languages presents both opportunities and challenges for developing multilingual code LLMs. While existing techniques often train code LLMs by simply aggregating multilingual code data, few explore the deeper relationships between programming languages and how such relationships can be utilized to optimize the training and inference of code LLMs. In this work, we investigate two fundamental questions: (1) What are the deep linguistic relationships among programming languages? and (2) How can these relationships be leveraged to improve multilingual code LLMs? We propose an embedding-based framework to uncover the latent families of programming languages. Our approach begins by defining 21 primary linguistic features of programming languages, such as variable definition, control structures, and method declarations, and then employs LLMs to generate feature-aligned code samples across multiple languages. By embedding these semantically parallel code snippets from 19 languages, we construct a similarity matrix and perform hierarchical clustering to uncover inherent language relationships. Our analysis reveals clear hierarchical structures among programming languages. Closely related languages form well-defined clusters (e.g., C, C++, Java, and Swift group together), while Go exhibits as a central language with the highest cross-language similarity. Building on the uncovered language families, we propose three strategies to enhance multilingual LLM training: transfer learning across linguistically related languages, linguistic proximity-guided curriculum learning, and centroid-based intermediary code translation. Experiments on four code intelligence tasks demonstrate that our methods significantly improve multilingual LLM performance. This work offers a universal perspective on programming languages and advances more effective strategies for multilingual code LLM training.
CCS Concepts: • Software and its engineering → Software maintenance and tools; • Computing methodologies → Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 被引用 378 次
- Cross-Domain Deep Code Search with Meta LearningYitian Chai, Hongyu Zhang, Beijun Shen, Xiaodong GuICSE 2022 · 被引用 42 次
- Syntax and Domain Aware Model for Unsupervised Program TranslationFang Liu, Jia Li, Li ZhangICSE 2023 · 被引用 25 次
- MetaTPTrans: A Meta Learning Approach for Multilingual Code Representation LearningWeiguo Pian, Hanyu Peng, Xunzhu Tang, Tiezhu Sun 等AAAI 2023 · 被引用 18 次
相关 Paper
- The Struggles of LLMs in Cross-Lingual Code Clone DetectionMicheline Bénédicte Moumoula, Abdoul Kader Kaboré, Jacques Klein, Tegawendé F. BissyandéFSE 2025 · 被引用 4 次
- Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?Zhe Yin, Xiaodong Gu, Beijun ShenFSE 2026
- IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code GeneratorsIndraneil Paul, Goran Glavas, Iryna GurevychACL 2024
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
- Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction TuningJian Yang, Wei Zhang, Yibo Miao, Shanghaoran Quan 等ACL 2025 · 被引用 4 次
