ConTested: Consistency-Aided Tested Code Generation with LLM
Jinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong, Dan Hao
Abstract
Recent advancements in large language models (LLMs) have significantly improved code generation, which generates code snippets automatically based on natural language requirements. Despite achieving state-of-the-art performance, LLMs often struggle to generate accurate and reliable code, requiring developers to spend substantial effort debugging and evaluating the generated output. Researchers have proposed leveraging Consistency to select code that passes more tests (inter-consistency) and demonstrates consistent behavior across more counterparts (intra-consistency). However, since the tests themselves are also generated by LLMs, relying on majority voting based on incorrect tests leads to unreliable results. To address this, we propose a lightweight interaction framework that incorporates user feedback to effectively guide consistency. Our results demonstrate that, with minimal human effort, performance can be significantly improved. In each iteration, we introduce a rank-correct-fix co-evolution process between code and tests. This process iteratively enhances the quality of both, making the consistency voting between code and tests more reliable. We evaluate ConTested through extensive experiments, demonstrating its effectiveness across multiple LLMs, including GPT-3.5 and GPT-4o. Our results show improvements of 32.9% over GPT-3.5 and 16.97% over GPT-4o. Additionally, ConTested achieves an 11.1% improvement over the SOTA post-processing technique, MPSC. This improvement is achieved with only a 4-round interaction with users, requiring minimal user effort. A user study further confirms the feasibility and cost-effectiveness of ConTested, highlighting its ability to enhance code generation without introducing substantial overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fddf418f-50ba-492a-a7ae-c956a2d03cc3Cited by top-tier papers2
- MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied AgentsGengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai WangOOPSLA 2026
- Validating LLM-Generated SQL Queries through Metamorphic PromptingLi Lin, Qinglin Zhu, Jintai Hong, Chong Wang et al.FSE 2026
Related papers
- Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationHaoxiang Jia, Robbie Morris, He Ye, Federica Sarro et al.ASE 2025 · 6 citations
- Enhancing Large Language Models in Coding Through Multi-Perspective Self-ConsistencyBaizhou Huang, Shuai Lu, Xiaojun Wan, Nan DuanACL 2024 · 4 citations
- Planning-Driven Programming: A Large Language Model Programming WorkflowChao Lei, Yanchuan Chang, Nir Lipovetzky, Krista A. EhingerACL 2025
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu et al.FSE 2024 · 49 citations
- Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language ModelsYanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen et al.FSE 2025 · 6 citations
