ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming
Xinwei Yang, Zhaofeng Liu, Chen Huang, Jiashuai Zhang, Tong Zhang, Yifan Zhang, Wenqiang Lei
摘要
While recent research increasingly emphasizes the value of human-LLM collaboration in competitive programming and proposes numerous empirical methods, a comprehensive understanding remains elusive due to the fragmented nature of existing studies and their use of diverse, application-specific human feedback. Thus, our work serves a three-fold purpose: First, we present the first taxonomy of human feedback consolidating the entire programming process, which promotes fine-grained evaluation. Second, we introduce ELABORA-TIONSET, a novel programming dataset specifically designed for human-LLM collaboration, meticulously annotated to enable large-scale simulated human feedback and facilitate costeffective real human interaction studies. Third, we introduce ELABORATION, a novel benchmark to facilitate a thorough assessment of human-LLM competitive programming. With ELABORATION, we pinpoint strengthes and weaknesses of existing methods, thereby setting the foundation for future improvement. Our code and dataset are available at https: //github.com/SCUNLP/ELABORATION .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AutoCode: LLMs as Problem Setters for Competitive ProgrammingShang Zhou, Zihan Zheng, Kaiyuan Liu, Zeyu Shen 等ICLR 2026 · 被引用 12 次
- E2EDev: Benchmarking Large Language Models in End-to-End Software Development TaskJingyao Liu, Chen Huang, Zhizhao Guan, Wenqiang Lei 等ACL 2026 · 被引用 5 次
- SPIRAL: Symbolic LLM Planning via Grounded and Reflective SearchYifan Zhang, Giridhar Ganapavarapu, Srideepika Jayaraman, Bhavna Agrawal 等AAAI 2026 · 被引用 4 次
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu 等ICML 2026 · 被引用 1 次
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao 等ICML 2026
相关 Paper
- RECODE-H: A Benchmark for Research Code Development with Interactive Human FeedbackChunyu Miao, Henry Peng Zou, Yangning Li, Yankai Chen 等ICLR 2026 · 被引用 25 次
- AetherCode: Evaluating LLMs’ Ability to Win In Premier Programming CompetitionsZihan Wang, Jiaze Chen, Zhicheng Liu, Haojie Pan 等ICLR 2026 · 被引用 11 次
- ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback EnvironmentsHojae Han, Seung-won Hwang, Rajhans Samdani, Yuxiong HeICLR 2025
- Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing FeedbackHannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella LapataACL 2025 · 被引用 7 次
- DyCodeEval: Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data ContaminationSimin Chen, Pranav Pusarla, Baishakhi RayICML 2025
