How Effectively Do Code Language Models Understand Poor-Readability Code?
Chao Hu, Yitian Chai, Hao Zhou, Fandong Meng, Jie Zhou, Xiaodong Gu
摘要
Code language models such as CodeT5 and CodeLlama have demonstrated substantial achievement in code comprehension. While the majority of research efforts have focused on improving model architectures and training processes, we find that the current benchmarks used for evaluating code comprehension models are confined to high-readability code, regardless of the popularity of low-readability code in reality. As such, they are inadequate to demonstrate the full spectrum of the model's ability, particularly the robustness to varying readability degrees. In this paper, we analyze the robustness of code summarization models to code with varying readability, including seven obfuscated datasets derived from existing benchmarks. Our findings indicate that current code summarization models are vulnerable to code with poor readability. In particular, their performance predominantly depends on semantic cues within the code, often neglecting the syntactic aspects. Existing benchmarks are biased toward evaluating semantic features, thereby overlooking the models' ability to understand nonsensitive syntactic features. Based on the findings, we present Poor-CodeSumEval, a new evaluation benchmark on code summarization tasks. PoorCodeSumEval innovatively introduces readability into the testing process, considering semantic, syntactic, and their cross-obfuscation, thereby providing a more comprehensive and rigorous evaluation of code summarization models. Our studies also provide more insightful suggestions for future research, such as constructing multi-readability benchmarks to evaluate the robustness of models on poor-readability code, proposing readability-awareness metrics, and automatic methods for code data cleaning and normalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion ModelsHaocheng Huang, Yuchen Chen, Weisong Sun, Peizhuo Lv 等FSE 2026
- What Makes In-Context Examples Effective for Code Generation?Dongze Li, Songqiang Chen, Jialun Cao, Shing-Chi CheungISSTA 2026
- TrapHunter: Exposing Covert Pathways in Trap Token ContractsYin Wu, Yixuan Liu, Yi Li, Chenyang Peng 等ISSTA 2026
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM BudgetDangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo 等ICSE 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 被引用 162 次
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev 等CCS 2018 · 被引用 148 次
相关 Paper
- Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the FamiliarYuanliang Zhang, Yifan Xie, Shanshan Li, Ke Liu 等ICSE 2025 · 被引用 2 次
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen 等ICSE 2022 · 被引用 76 次
- Hallucinations in LLM-Based Code Summarization: Unveiling, Detection, and MitigationGuanghua Wan, Yuanning Feng, Yao Wan, Zhaoyang Chu 等FSE 2026
- ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation GroundingIndraneil Paul, Haoyi Yang, Goran Glavas, Kristian Kersting 等ICLR 2025
- ReCode: Robustness Evaluation of Code Generation ModelsShiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang 等ACL 2023 · 被引用 32 次
