DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code Generation
Jizheng Chen, Kounianhua Du, Xinyi Dai, Weiming Zhang, Xihuai Wang, Yasheng Wang, Ruiming Tang, Weinan Zhang, Yong Yu
摘要
With the impressive reasoning and text generation capabilities of large language models (LLMs), methods leveraging multiple LLMs to debate each other have garnered increasing attention. However, existing debate-based approaches remain limited in effectiveness in structured and detailed domains represented by code generation due to several reasons: 1) Reliance on different instances of the same LLM for debate, neglecting the potential benefits of integrating diverse models with varied internal knowledge for more comprehensive code generation, 2) under-utilization of test cases, and 3) reliance on third-party LLM moderators for result consolidation and decision-making, probably introducing hallucinations and judgment errors. To address these challenges, we propose DebateCoder to collect intelligence of LLMs via test case-driven debate for code generation. In DebateCoder, test cases serve as a medium for models to analyze code and identify bugs, while opposing models generate test cases to challenge each other's code during the debate process. These test cases, along with their execution results, are elaborately leveraged to refine and enhance the code through a novel contrastive analysis process. Furthermore, Debate-Coder leverages test case outcomes to assess code quality and determine convergence criteria. Unlike previous approaches, DebateCoder emphasizes the collaborative improvement of both models through competitive debate and interactive analysis. Abundant experimental results on two datasets demonstrate the effectiveness of DebateCoder.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code DebuggingWeiming Zhang, Qingyao Li, Xinyi Dai, Jizheng Chen 等EMNLP 2025 · 被引用 1 次
- Natural Language-Focused Software Engineering via Code-Documentation EquivalenceAryaz Eghbali, Zhongxin Liu, Michael PradelFSE 2026
它引用的顶会 Paper19
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
相关 Paper
- A Pair Programming Framework for Code Generation via Multi-Plan Exploration and Feedback-Driven RefinementHuan Zhang, Wei Cheng, Yuhan Wu, Wei HuASE 2024 · 被引用 7 次
- BiVCoder: A Multi-Agent Framework for Code Generation via Bidirectional Code-Test DiagnosisXiaoyang Li, Jinhao Dong, Wenhang Shi, Wei Lu 等KDD 2026
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li 等ICLR 2026 · 被引用 4 次
- CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsJingwei Shi, Xinxiang Yin, Jing Huang, Shengyu Tao 等ACL 2026 · 被引用 6 次
- Revisit Self-Debugging with Self-Generated Tests for Code GenerationXiancai Chen, Zhengwei Tao, Kechi Zhang, Changzhi Zhou 等ACL 2025
