ConTested: Consistency-Aided Tested Code Generation with LLM
Jinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong, Dan Hao
摘要
Recent advancements in large language models (LLMs) have significantly improved code generation, which generates code snippets automatically based on natural language requirements. Despite achieving state-of-the-art performance, LLMs often struggle to generate accurate and reliable code, requiring developers to spend substantial effort debugging and evaluating the generated output. Researchers have proposed leveraging Consistency to select code that passes more tests (inter-consistency) and demonstrates consistent behavior across more counterparts (intra-consistency). However, since the tests themselves are also generated by LLMs, relying on majority voting based on incorrect tests leads to unreliable results. To address this, we propose a lightweight interaction framework that incorporates user feedback to effectively guide consistency. Our results demonstrate that, with minimal human effort, performance can be significantly improved. In each iteration, we introduce a rank-correct-fix co-evolution process between code and tests. This process iteratively enhances the quality of both, making the consistency voting between code and tests more reliable. We evaluate ConTested through extensive experiments, demonstrating its effectiveness across multiple LLMs, including GPT-3.5 and GPT-4o. Our results show improvements of 32.9% over GPT-3.5 and 16.97% over GPT-4o. Additionally, ConTested achieves an 11.1% improvement over the SOTA post-processing technique, MPSC. This improvement is achieved with only a 4-round interaction with users, requiring minimal user effort. A user study further confirms the feasibility and cost-effectiveness of ConTested, highlighting its ability to enhance code generation without introducing substantial overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied AgentsGengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai WangOOPSLA 2026
- Validating LLM-Generated SQL Queries through Metamorphic PromptingLi Lin, Qinglin Zhu, Jintai Hong, Chong Wang 等FSE 2026
相关 Paper
- Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationHaoxiang Jia, Robbie Morris, He Ye, Federica Sarro 等ASE 2025 · 被引用 6 次
- Enhancing Large Language Models in Coding Through Multi-Perspective Self-ConsistencyBaizhou Huang, Shuai Lu, Xiaojun Wan, Nan DuanACL 2024 · 被引用 4 次
- Planning-Driven Programming: A Large Language Model Programming WorkflowChao Lei, Yanchuan Chang, Nir Lipovetzky, Krista A. EhingerACL 2025
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu 等FSE 2024 · 被引用 49 次
- Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language ModelsYanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen 等FSE 2025 · 被引用 6 次
