CrystalBLEU: Precisely and Efficiently Measuring the Similarity of Code
Aryaz Eghbali, Michael Pradel
摘要
Recent years have brought a surge of work on predicting pieces of source code, e.g., for code completion, code migration, program repair, or translating natural language into code. All this work faces the challenge of evaluating the quality of a prediction w.r.t. some oracle, typically in the form of a reference solution. A common evaluation metric is the BLEU score, an n-gram-based metric originally proposed for evaluating natural language translation, but adopted in software engineering because it can be easily computed on any programming language and enables automated evaluation at scale. However, a key difference between natural and programming languages is that in the latter, completely unrelated pieces of code may have many common n-grams simply because of the syntactic verbosity and coding conventions of programming languages. We observe that these trivially shared n-grams hamper the ability of the metric to distinguish between truly similar code examples and code examples that are merely written in the same language. This paper presents CrystalBLEU, an evaluation metric based on BLEU, that allows for precisely and efficiently measuring the similarity of code. Our metric preserves the desirable properties of BLEU, such as being language-agnostic, able to handle incomplete or partially incorrect code, and efficient, while reducing the noise caused by trivially shared n-grams. We evaluate CrystalBLEU on two datasets from prior work and on a new, labeled dataset of semantically equivalent programs. Our results show that CrystalBLEU can distinguish similar from dissimilar code examples 1.9-4.5 times more effectively, when compared to the original BLEU score and a previously proposed variant of BLEU for code. CCS CONCEPTS • General and reference → Metrics; Evaluation; • Software and its engineering → Software notations and tools; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- OctoPack: Instruction Tuning Code Large Language ModelsNiklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng 等ICLR 2024 · 被引用 203 次
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang 等ICSE 2024 · 被引用 124 次
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
- CodeBERTScore: Evaluating Code Generation with Pretrained Models of CodeShuyan Zhou, Uri Alon, Sumit Agarwal, Graham NeubigEMNLP 2023 · 被引用 77 次
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 被引用 64 次
它引用的顶会 Paper12
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2020 · 被引用 242 次
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
相关 Paper
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 被引用 8 次
- Evaluating Code Summarization Techniques: A New Metric and an Empirical CharacterizationAntonio Mastropaolo, Matteo Ciniselli, Massimiliano Di Penta, Gabriele BavotaICSE 2024 · 被引用 28 次
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 被引用 103 次
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen 等ICSE 2022 · 被引用 76 次
