SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
Qinglin Wang, Zhihong Sun, Ruyun Wang, Tao Huang, Zhi Jin, Ge Li, Chen Lyu
Abstract
Large Language Models (LLMs) can translate natural language requirements into code, yet empirical analyses of representative models reveal that semantic errors—programs that compile but behave incorrectly—constitute the majority of observed faults (e.g., >60% on DeepSeek-Coder-6.7B and QwenCoder-7B). Post-hoc repair pipelines detect such faults only after execution, incurring latency, relying on incomplete test suites, and often mis-localizing the defect. Since semantic drift originates in the autoregressive decoding process, intervening while the code is being generated is a direct way to stop error propagation. Constrained-decoding approaches such as ROCODE attempt this, but still wait until the entire program runs to obtain feedback and use entropy heuristics that do not truly capture semantics. A more effective solution must inject semantic signals—early and precisely—into the decoding process. We present SemGuard, a semantic-evaluator-driven framework that performs real-time, line-level semantic supervision. To train the evaluator, we build SemDiff, the first dataset with fine-grained annotations that mark the exact line where a correct and an incorrect implementation diverge. The evaluator, once embedded in the LLM’s decoder, flags deviations on partial code, rolls back to the faulty line, and guides regeneration—without executing the program or requiring test cases. Across four benchmarks, SemGuard consistently outperforms state-of-the-art baselines. It lowers the semantic error rate by 19.86% on SemDiff relative to ROCODE, and lifts Pass@1 by 48.92% on the realworld LiveCodeBench with CodeLlama-7B. Similar gains hold for StarCoder2-7B on MBPP and for DeepSeekCoder-6.7B on the Java benchmark SemDiff-Java, demonstrating model- and language-agnostic effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 98 citations
- Hot or Cold? Adaptive Temperature Sampling for Code Generation with Large Language ModelsYuqi Zhu, Jia Li, Ge Li, Yunfei Zhao et al.AAAI 2024 · 68 citations
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan et al.ICLR 2023 · 64 citations
Related papers
- ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code GenerationXue Jiang, Yihong Dong, Yongding Tao, Huanyu Liu et al.ICSE 2025 · 6 citations
- TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code GenerationHenrijs Princis, Arindam Sharma, Cristina DavidPLDI 2026
- SAFENUDGE: Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offsJoão Fonseca, Andrew Bell, Julia StoyanovichEMNLP 2025 · 1 citation
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu et al.ISSTA 2024 · 10 citations
- Hallucinations in LLM-Based Code Summarization: Unveiling, Detection, and MitigationGuanghua Wan, Yuanning Feng, Yao Wan, Zhaoyang Chu et al.FSE 2026
