ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming
Xinwei Yang, Zhaofeng Liu, Chen Huang, Jiashuai Zhang, Tong Zhang, Yifan Zhang, Wenqiang Lei
Abstract
While recent research increasingly emphasizes the value of human-LLM collaboration in competitive programming and proposes numerous empirical methods, a comprehensive understanding remains elusive due to the fragmented nature of existing studies and their use of diverse, application-specific human feedback. Thus, our work serves a three-fold purpose: First, we present the first taxonomy of human feedback consolidating the entire programming process, which promotes fine-grained evaluation. Second, we introduce ELABORA-TIONSET, a novel programming dataset specifically designed for human-LLM collaboration, meticulously annotated to enable large-scale simulated human feedback and facilitate costeffective real human interaction studies. Third, we introduce ELABORATION, a novel benchmark to facilitate a thorough assessment of human-LLM competitive programming. With ELABORATION, we pinpoint strengthes and weaknesses of existing methods, thereby setting the foundation for future improvement. Our code and dataset are available at https: //github.com/SCUNLP/ELABORATION .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- AutoCode: LLMs as Problem Setters for Competitive ProgrammingShang Zhou, Zihan Zheng, Kaiyuan Liu, Zeyu Shen et al.ICLR 2026 · 12 citations
- E2EDev: Benchmarking Large Language Models in End-to-End Software Development TaskJingyao Liu, Chen Huang, Zhizhao Guan, Wenqiang Lei et al.ACL 2026 · 5 citations
- SPIRAL: Symbolic LLM Planning via Grounded and Reflective SearchYifan Zhang, Giridhar Ganapavarapu, Srideepika Jayaraman, Bhavna Agrawal et al.AAAI 2026 · 4 citations
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu et al.ICML 2026 · 1 citation
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao et al.ICML 2026
Related papers
- RECODE-H: A Benchmark for Research Code Development with Interactive Human FeedbackChunyu Miao, Henry Peng Zou, Yangning Li, Yankai Chen et al.ICLR 2026 · 25 citations
- AetherCode: Evaluating LLMs’ Ability to Win In Premier Programming CompetitionsZihan Wang, Jiaze Chen, Zhicheng Liu, Haojie Pan et al.ICLR 2026 · 11 citations
- ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback EnvironmentsHojae Han, Seung-won Hwang, Rajhans Samdani, Yuxiong HeICLR 2025
- Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing FeedbackHannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella LapataACL 2025 · 7 citations
- DyCodeEval: Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data ContaminationSimin Chen, Pranav Pusarla, Baishakhi RayICML 2025
