CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
Monoshi Kumar Roy, Simin Chen, Benjamin Steenhoek, Jinjun Peng, Gail E. Kaiser, Baishakhi Ray, Wei Le
摘要
Understanding and reasoning about code semantics is essential for enhancing code LLMs' abilities to solve real-world software engineering (SE) tasks. Although several code reasoning benchmarks exist, most rely on synthetic datasets or educational coding problems and focus on coarse-grained reasoning tasks such as input/output prediction, limiting their effectiveness in evaluating LLMs in practical SE contexts. To bridge this gap, we propose CodeSense, the first benchmark that makes available a spectrum of fine-grained code reasoning tasks concerned with the software engineering of real-world code. We collected Python, C and Java software projects from real-world repositories. We executed tests from these repositories, collected their execution traces, and constructed a ground truth dataset for fine-grained semantic reasoning tasks. We then performed comprehensive evaluations on state-of-the-art LLMs. Our results show a clear performance gap for the models to handle fine-grained reasoning tasks. Although prompting techniques such as chain-of-thought and in-context learning helped, the lack of code semantics in LLMs fundamentally limit models' capabilities of code reasoning. Besides dataset, benchmark and evaluation, our work produced an execution tracing framework and tool set that make it easy to collect ground truth for fine-grained SE reasoning tasks, offering a strong basis for future benchmark construction and model post training. Our code and data are located at https://codesense-bench.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Is “Knowing It’s Malicious” Enough? Evaluating LLMs for Fine-Grained Malware Behavior AuditingXinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang 等ISSTA 2026
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language ModelsChangshu Liu, Yang Chen, Reyhaneh JabbarvandICSE 2026
- StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHao Wang, Lei Sha, Jie ZhangICML 2026
- RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language ModelsYanlin Wang, Suiquan Wang, Yanli Wang, Bowen Zhang 等FSE 2026
它引用的顶会 Paper14
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama 等ICML 2024 · 被引用 270 次
- Self-Supervised Bug Detection and RepairMiltiadis Allamanis, Henry Jackson-Flux, Marc BrockschmidtNeurIPS 2021 · 被引用 145 次
相关 Paper
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 等ICLR 2025
- CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMsDung Manh Nguyen, Thang Chau Phan, Nam Le Hai, Tien-Thong Doan 等ICLR 2025
- LongCodeU: Benchmarking Long-Context Language Models on Long Code UnderstandingJia Li, Xuyuan Guo, Lei Li, Kechi Zhang 等ACL 2025
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 被引用 9 次
