CrystalBLEU: Precisely and Efficiently Measuring the Similarity of Code
Aryaz Eghbali, Michael Pradel
Abstract
Recent years have brought a surge of work on predicting pieces of source code, e.g., for code completion, code migration, program repair, or translating natural language into code. All this work faces the challenge of evaluating the quality of a prediction w.r.t. some oracle, typically in the form of a reference solution. A common evaluation metric is the BLEU score, an n-gram-based metric originally proposed for evaluating natural language translation, but adopted in software engineering because it can be easily computed on any programming language and enables automated evaluation at scale. However, a key difference between natural and programming languages is that in the latter, completely unrelated pieces of code may have many common n-grams simply because of the syntactic verbosity and coding conventions of programming languages. We observe that these trivially shared n-grams hamper the ability of the metric to distinguish between truly similar code examples and code examples that are merely written in the same language. This paper presents CrystalBLEU, an evaluation metric based on BLEU, that allows for precisely and efficiently measuring the similarity of code. Our metric preserves the desirable properties of BLEU, such as being language-agnostic, able to handle incomplete or partially incorrect code, and efficient, while reducing the noise caused by trivially shared n-grams. We evaluate CrystalBLEU on two datasets from prior work and on a new, labeled dataset of semantically equivalent programs. Our results show that CrystalBLEU can distinguish similar from dissimilar code examples 1.9-4.5 times more effectively, when compared to the original BLEU score and a previously proposed variant of BLEU for code. CCS CONCEPTS • General and reference → Metrics; Evaluation; • Software and its engineering → Software notations and tools; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc6d7487-78b2-449b-a445-ff9363ad34b2Cited by top-tier papers26
- OctoPack: Instruction Tuning Code Large Language ModelsNiklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng et al.ICLR 2024 · 203 citations
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang et al.ICSE 2024 · 124 citations
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- CodeBERTScore: Evaluating Code Generation with Pretrained Models of CodeShuyan Zhou, Uri Alon, Sumit Agarwal, Graham NeubigEMNLP 2023 · 77 citations
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 64 citations
Builds on12
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik et al.ICLR 2020 · 212 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
Related papers
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 8 citations
- Evaluating Code Summarization Techniques: A New Metric and an Empirical CharacterizationAntonio Mastropaolo, Matteo Ciniselli, Massimiliano Di Penta, Gabriele BavotaICSE 2024 · 28 citations
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 103 citations
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen et al.ICSE 2022 · 76 citations
