Dependency-Aware Code Naturalness
Chen Yang, Junjie Chen, Jiajun Jiang, Yuliang Huang
摘要
Code naturalness, which captures repetitiveness and predictability in programming languages, has proven valuable for various code-related tasks in software engineering. However, precisely measuring code naturalness remains a fundamental challenge. Existing methods measure code naturalness over individual lines of code while ignoring the deep semantic relations among different lines, e.g., program dependency, which may negatively affect the precision of the measure. Despite the intuitive appeal of extending the code naturalness measure to the code dependency domain (as there are some work that have initiated the utilization of code dependency for diverse code-related tasks), this assumption remains unexplored and warrants direct investigation. In this study, we aim to perform the first empirical study to investigate whether incorporating code dependency, instead of analyzing individual lines, can enhance the precision of measuring code naturalness.
To achieve that, we first propose a new method named DAN for measuring code naturalness by incorporating the rich dependency information in the code. Specifically, DAN extracts multiple sequences of code lines by traversing the program dependency graph, where different code lines are connected by dependencies in each sequence, and then the code naturalness will be measured by taking each sequence as a whole. In this way, the dependency information can be well captured. Finally, we have conducted an extensive study to evaluate the influence of code dependency for measuring code naturalness with DAN, and compared it with the state-of-the-art methods under three emerging application scenarios of code naturalness. The results demonstrate that DAN can not only better distinguish natural and unnatural code, but also substantially boost two important downstream applications of code naturalness, i.e., distinguishing buggy and non-buggy code lines and data cleansing for training better code models, reflecting the significance of code dependency in measuring code naturalness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fuzzing MLIR Compiler Infrastructure via Operation Dependency AnalysisChenyao Suo, Junjie Chen, Shuang Liu, Jiajun Jiang 等ISSTA 2024 · 被引用 12 次
- Reflective Unit Test Generation for Precise Type Error Detection with Large Language ModelsChen Yang, Ziqi Wang, Yanjie Jiang, Lin Yang 等ASE 2025 · 被引用 1 次
- Clarifying Semantics of In-Context Examples for Unit Test GenerationChen Yang, Lin Yang, Ziqi Wang, Dong Wang 等ASE 2025 · 被引用 1 次
- Uncovering Business Logic Bugs via Semantics-Driven Unit Test Generation (Experience Paper)Chen Yang, Junjie ChenISSTA 2026
它引用的顶会 Paper15
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- Is Self-Repair a Silver Bullet for Code Generation?Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao 等ICLR 2024 · 被引用 195 次
相关 Paper
- Do bugs lead to unnaturalness of source code?Yanjie Jiang, Hui Liu, Yuxia Zhang, Weixing Ji 等FSE 2022 · 被引用 8 次
- Has My Code Been Stolen for Model Training? A Naturalness Based Approach to Code Contamination DetectionHaris Ali Khan, Yanjie Jiang, Qasim Umer, Yuxia Zhang 等FSE 2025 · 被引用 1 次
- Beyond Sequences: Two-dimensional Representation and Dependency Encoding for Code GenerationXiangyu Zhang, Yu Zhou, Guang Yang, Wei Cheng 等ACL 2025
- Contextuality of Code Representation LearningYi Li, Shaohua Wang, Tien N. NguyenASE 2023 · 被引用 2 次
- CrystalBLEU: Precisely and Efficiently Measuring the Similarity of CodeAryaz Eghbali, Michael PradelASE 2022 · 被引用 33 次
