The Devil is in the Tails: How Long-Tailed Code Distributions Impact Large Language Models
Xin Zhou, Kisub Kim, Bowen Xu, Jiakun Liu, DongGyun Han, David Lo
Abstract
Learning-based techniques, especially advanced Large Language Models (LLMs) for code, have gained considerable popularity in various software engineering (SE) tasks. However, most existing works focus on designing better learning-based models and pay less attention to the properties of datasets. Learning-based models, including popular LLMs for code, heavily rely on data, and the data's properties (e.g., data distribution) could significantly affect their behavior. We conducted an exploratory study on the distribution of SE data and found that such data usually follows a skewed distribution (i.e., long-tailed distribution) where a small number of classes have an extensive collection of samples, while a large number of classes have very few samples. We investigate three distinct SE tasks and analyze the impacts of long-tailed distribution on the performance of LLMs for code. Our experimental results reveal that the long-tailed distribution has a substantial impact on the effectiveness of LLMs for code. Specifically, LLMs for code perform between 30.0% and 254.0% worse on data samples associated with infrequent labels compared to data samples of frequent labels. Our study provides a better understanding of the effects of long-tailed distributions on popular LLMs for code and insights for the future development of SE automation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88f41ba2-dfb7-4d2d-8808-9ec1525d3f8bCited by top-tier papers7
- Out of Sight, Out of Mind: Better Automatic Vulnerability Repair by Broadening Input Ranges and SourcesXin Zhou, Kisub Kim, Bowen Xu, DongGyun Han et al.ICSE 2024 · 32 citations
- Large Language Models for In-File Vulnerability Localization Can Be "Lost in the End"Francesco Sovrano, Adam Bauer, Alberto BacchelliFSE 2025 · 8 citations
- One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)Xu Yang, Shaowei Wang, Jiayuan Zhou, Wenhan ZhuFSE 2025 · 7 citations
- Applying Contrastive Learning to Code Vulnerability Type ClassificationChen Ji, Su Yang, Hongyu Sun, Yuqing ZhangEMNLP 2024 · 6 citations
- Instructive Code Retriever: Learn from Large Language Model's Feedback for Code Intelligence TasksJiawei Lu, Haoye Wang, Zhongxin Liu, Keyu Liang et al.ASE 2024 · 3 citations
Builds on18
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 150 citations
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella et al.ICSE 2022 · 149 citations
Related papers
- Scaling Laws for Code: A More Data-Hungry RegimeXianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang et al.ACL 2026 · 3 citations
- LLM-based Vulnerability Discovery through the Lens of Code MetricsFelix Weissberg, Lukas Pirch, Erik Imgrund, Jonas Möller et al.ICSE 2026
- Evaluating Large Language Models in Class-Level Code GenerationXueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang et al.ICSE 2024 · 118 citations
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang et al.NeurIPS 2024 · 26 citations
- Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution DetectionYanfu Yan, Viet Duong, Huajie Shao, Denys PoshyvanykICSE 2025 · 1 citation
