Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors
Bohan Lyu, Siqiao Huang, Zichen Liang, Qian Sun, Jiaming Zhang
摘要
Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks, such as generation and understanding. However, an equally important yet underexplored question is whether LLMs can serve as surrogate models for code execution prediction. To systematically investigate it, we introduce SURGE, a comprehensive benchmark with 1160 problems covering 8 key aspects: multilanguage programming tasks, competitionlevel programming problems, repository-level code analysis, high-cost scientific computing, time-complexity-intensive algorithms, buggy code analysis, programs dependent on specific compilers or execution environments, and formal mathematical proof verification. Through extensive analysis of 21 open-source and proprietary LLMs, we examine scaling laws, data efficiency, and predictive accuracy. Our findings reveal important insights about the feasibility of LLMs as efficient surrogates for computational processes. The benchmark and evaluation framework are available at https: //github.com/Imbernoulli/SURGE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li 等ISSTA 2020 · 被引用 325 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
- LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific DiscoveryPingchuan Ma, Tsun-Hsuan Wang, Minghao Guo, Zhiqing Sun 等ICML 2024 · 被引用 76 次
相关 Paper
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu 等ACL 2023 · 被引用 104 次
- CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationWeixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li 等ACL 2024
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang 等ACL 2024 · 被引用 21 次
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationZhihao Gong, Zeyu Sun, Dong Huang, Qingyuan Liang 等ACL 2026 · 被引用 5 次
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu 等ICML 2026 · 被引用 1 次
