PGS: Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback
Lehan He, Zeren Chen, Zhe Zhang, Xiang Gao, Lu Sheng
Abstract
Large Language Models (LLMs) excel at code generation, yet ensuring the functional correctness of their outputs remains a persistent challenge. Recent studies have applied Test-Driven Development (TDD) to refine code, leveraging execution feedback to guide the model toward correct solutions. However, such feedback is often noisy and uninformative, stemming from the scarcity of high-quality test cases and the abundance of noisy, auto-generated ones. In this work, we shift the focus from test-case generation to feedback quality. We introduce the Property-Generated Solver (PGS), a novel feedback-centric framework designed to generate highly effective feedback via two principles: it must provide semantic guidance beyond simple I/O mismatches through property validation, and be structurally minimal, to reduce cognitive load and isolate root causes. PGS operates by checking high-level program properties (e.g., a sorting function must produce a non-decreasing sequence) then providing the simplest failing counterexample to the LLM. This property-driven, minimal feedback steers LLMs toward correct and generalizable solutions. Across diverse benchmarks, PGS demonstrates superior performance, achieving a bug fix rate 1.4x-1.6x higher than the strongest debugging-based approaches and establishing a new state-of-the-art in automated code refinement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77a7ba1a-f178-4986-a80c-7839416ef4dcBuilds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
Related papers
- Revisit Self-Debugging with Self-Generated Tests for Code GenerationXiancai Chen, Zhengwei Tao, Kechi Zhang, Changzhi Zhou et al.ACL 2025
- Test-Driven Development and LLM-based Code GenerationNoble Saji Mathews, Meiyappan NagappanASE 2024 · 15 citations
- Learner-Tailored Program Repair: A Solution Generator with Iterative Edit-Driven Retrieval EnhancementZhenlong Dai, Zhuoluo Zhao, Hengning Wang, Xiu Tang et al.AAAI 2026
- SeDev: Structured Semantic Exploration for LLM-Driven Code GenerationRonghui Yang, Jie Liu, Jiajie Zeng, Jiexin Wang et al.ACL 2026
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit TestsJunda Zhao, Shurui Zhou, Eldan CohenISSTA 2026 · 1 citation
