CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
Monoshi Kumar Roy, Simin Chen, Benjamin Steenhoek, Jinjun Peng, Gail E. Kaiser, Baishakhi Ray, Wei Le
Abstract
Understanding and reasoning about code semantics is essential for enhancing code LLMs' abilities to solve real-world software engineering (SE) tasks. Although several code reasoning benchmarks exist, most rely on synthetic datasets or educational coding problems and focus on coarse-grained reasoning tasks such as input/output prediction, limiting their effectiveness in evaluating LLMs in practical SE contexts. To bridge this gap, we propose CodeSense, the first benchmark that makes available a spectrum of fine-grained code reasoning tasks concerned with the software engineering of real-world code. We collected Python, C and Java software projects from real-world repositories. We executed tests from these repositories, collected their execution traces, and constructed a ground truth dataset for fine-grained semantic reasoning tasks. We then performed comprehensive evaluations on state-of-the-art LLMs. Our results show a clear performance gap for the models to handle fine-grained reasoning tasks. Although prompting techniques such as chain-of-thought and in-context learning helped, the lack of code semantics in LLMs fundamentally limit models' capabilities of code reasoning. Besides dataset, benchmark and evaluation, our work produced an execution tracing framework and tool set that make it easy to collect ground truth for fine-grained SE reasoning tasks, offering a strong basis for future benchmark construction and model post training. Our code and data are located at https://codesense-bench.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd53ebd8-01db-4a44-9325-d9b996c732cbCited by top-tier papers4
- Is “Knowing It’s Malicious” Enough? Evaluating LLMs for Fine-Grained Malware Behavior AuditingXinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang et al.ISSTA 2026
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language ModelsChangshu Liu, Yang Chen, Reyhaneh JabbarvandICSE 2026
- StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHao Wang, Lei Sha, Jie ZhangICML 2026
- RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language ModelsYanlin Wang, Suiquan Wang, Yanli Wang, Bowen Zhang et al.FSE 2026
Builds on14
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- Self-Supervised Bug Detection and RepairMiltiadis Allamanis, Henry Jackson-Flux, Marc BrockschmidtNeurIPS 2021 · 145 citations
Related papers
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu et al.ICLR 2025
- CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMsDung Manh Nguyen, Thang Chau Phan, Nam Le Hai, Tien-Thong Doan et al.ICLR 2025
- LongCodeU: Benchmarking Long-Context Language Models on Long Code UnderstandingJia Li, Xuyuan Guo, Lei Li, Kechi Zhang et al.ACL 2025
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 9 citations
