AdapTrack: Constrained Decoding without Distorting LLM's Output Intent
Yongmin Li, Jia Li, Ge Li, Zhi Jin
摘要
Language model-based code generation and completion tools have been widely adopted, but they may sometimes produce code that does not meet necessary constraints, such as syntactic correctness or API existence. Constrained decoding techniques are developed to help the model generate code that adheres to the constraints by greedily eliminating generation options that violate constraints at each step of the generation process. However, there is a severe limitation of constrained decoding, that it distorts the model’s output intent, forcing it to produce code that may satisfy the constraint but does not match the development intent and is therefore incorrect. In response to this challenge, we propose AdapTrack. By incorporating backtracking into the generation process, AdapTrack avoids distorting the output intent of the language model, thereby producing results that are not only constraint-compliant but also more semantically aligned with model’s output intent. On our synthetic API completion dataset, AdapTrack can achieve up to 360.87% improvement compared to constrained decoding; on the real-world API completion dataset we collect that exhibits similar issues, AdapTrack can also achieve an improvement of up to 38.93% over constrained decoding; in general code genration benchmarks, compared to constrained decoding, AdapTrack can achieve up to 7.84% improvement on HumanEval, and up to 6.42% improvement on MBPP. This indicates that, simply by better adhering to the model’s output intent, AdapTrack can achieve significant improvements. We provide a theoretical proof that the distribution produced by AdapTrack aligns with the model’s distribution given the generated tokens, thereby ensuring that the model’s output intent is not distorted. Experiments on domain-specific language problems show that, compared to existing methods, our approach can provide generation results that are more consistent with the language model’s distribution. Our code and supplementary material are available at https://github.com/Gompyn/adaptrack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Synchromesh: Reliable Code Generation from Pre-trained Language ModelsGabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari 等ICLR 2022 · 被引用 200 次
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program RepairYuxiang Wei, Chunqiu Steven Xia, Lingming ZhangFSE 2023 · 被引用 111 次
相关 Paper
- ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code GenerationXue Jiang, Yihong Dong, Yongding Tao, Huanyu Liu 等ICSE 2025 · 被引用 6 次
- Grammar-Aligned DecodingKanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova 等NeurIPS 2024 · 被引用 73 次
- Type-Constrained Code Generation with Language ModelsNiels Mündler, Jingxuan He, Hao Wang, Koushik Sen 等PLDI 2025 · 被引用 10 次
- ReCode: Robustness Evaluation of Code Generation ModelsShiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang 等ACL 2023 · 被引用 32 次
- Constrained Decoding of Diffusion LLMs with Context-Free GrammarsNiels Mündler, Jasper Dekoninck, Martin VechevICLR 2026 · 被引用 18 次
