Evaluating Large Language Models in Class-Level Code Generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, Yiling Lou
摘要
Recently, many large language models (LLMs) have been proposed, showing advanced proficiency in code generation. Meanwhile, many efforts have been dedicated to evaluating LLMs on code generation benchmarks such as HumanEval. Although being very helpful for comparing different LLMs, existing evaluation focuses on a simple code generation scenario (i.e., function-level or statement-level code generation), which mainly asks LLMs to generate one single code unit (e.g., a function or a statement) for the given natural language description. Such evaluation focuses on generating independent and often small-scale code units, thus leaving it unclear how LLMs perform in real-world software development scenarios. To fill this knowledge gap, we make the first attempt to evaluate LLMs in a more challenging code generation scenario, i.e., classlevel code generation. Compared with existing code generation benchmarks, it better reflects real-world software development scenarios due to it comprising broader contextual dependencies and multiple, interdependent units of code. We first manually construct the first class-level code generation benchmark ClassEval of 100 class-level Python code generation tasks with approximately 500 person-hours. Based on the new benchmark ClassEval, we then perform the first study of 11 state-of-the-art LLMs on class-level code generation. Based on our results, we find that all LLMs perform much worse on class-level code generation compared to the methodlevel. While GPT models still dominate other LLMs on class-level code generation, the performance rankings of other models on method-level code generation no longer holds for class-level code generation. Besides, most models (except GPT models) perform better when generating the class method by method; and they have the limited ability of generating dependent code. Based on our findings, we call for software engineering (SE) researchers' expertise to build more LLM benchmarks based on practical and complicated software development scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang 等ASE 2024 · 被引用 42 次
- CRUXEVAL-X: A Benchmark for Multilingual Code Reasoning, Understanding and ExecutionRuiyang Xu, Jialun Cao, Yaojie Lu, Ming Wen 等ACL 2025 · 被引用 26 次
- CODERL+: Improving Code Generation via Reinforcement with Execution Semantics AlignmentXue Jiang, Yihong Dong, Mengyang Liu, Hongyi Deng 等ACL 2026 · 被引用 18 次
- Test-Driven Development and LLM-based Code GenerationNoble Saji Mathews, Meiyappan NagappanASE 2024 · 被引用 15 次
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationDominik Fuchß, Tobias Hey, Jan Keim, Haoyu Liu 等ICSE 2025 · 被引用 8 次
它引用的顶会 Paper22
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 等NeurIPS 2023 · 被引用 810 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
相关 Paper
- ClassEval-T: Evaluating Large Language Models in Class-Level Code TranslationPengyu Xue, Linhao Wu, Zhen Yang, Chengyi Wang 等ISSTA 2025 · 被引用 5 次
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin 等AAAI 2025 · 被引用 25 次
- HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code GenerationDewu Zheng, Yanlin Wang, Ensheng Shi, Ruikai Zhang 等ICSE 2025 · 被引用 2 次
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong 等ASE 2024 · 被引用 7 次
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang 等FSE 2026
