Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs
Hantao Hua, Jiming Su, hao tang, Yiping Yao, Feng Zhu
Abstract
Grammar-constrained decoding enables large language models (LLMs) to reliably generate structured outputs such as JSON, SQL, and domain-specific programs. Existing systems often enforce constraints by executing byte-level parser logic inside the token-level decoding loop, introducing CPU-side control flow and CPU--GPU synchronization that become bottlenecks under continuous batching. We propose Gram2Token, a GPU-native framework that preprocesses deterministic byte-level grammar execution into token-level transitions before inference. Gram2Token aligns tokenizer byte sequences with grammar transitions through a trie and groups tokens with identical transition outcomes across preprocessed grammar states. These categories yield compact validity masks and transition tables, reducing run-time enforcement to category lookup, masking, and state update rather than parser-style byte traversal. Across four model families under schema-diverse continuous batching, Gram2Token achieves a geometric-mean throughput improvement of 1.38 over the strongest baseline, with a maximum speedup of 1.85, at the cost of additional preprocessing and time-to-first-token overhead. Break-even and grammar-complexity analyses show that this overhead is amortized by grammar reuse, longer outputs, and larger batches. These results show that token-level grammar preprocessing is an effective design point for high-throughput structured LLM serving. Code is available at https://github.com/Paradozile/Gram2Token.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4aa8db81-bd13-4412-8082-596ac57b3a0aBuilds on11
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Large Language Models as Tool MakersTianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen et al.ICLR 2024 · 283 citations
- Grammar-Aligned DecodingKanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova et al.NeurIPS 2024 · 73 citations
Related papers
- Efficient Grammar-Constrained Decoding via Parser Stack ClassificationYongmin Li, Yihong Dong, Jia Li, Ge LiISSTA 2026
- Flexible and Efficient Grammar-Constrained DecodingKanghee Park, Timothy Zhou, Loris D'AntoniICML 2025
- Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM GenerationJunyi Chen, Shihao Bai, Zaijun Wang, Siyu Wu et al.ACL 2025
- Earley-Driven Dynamic Pruning for Efficient Structured DecodingXintong Sun, Chi Wei, Minghao Tian, Shiwen NiICML 2025
- The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained DecodingAvinash Reddy, Thayne Walker, Jaime Ide, Amrit Singh BediICML 2026 · 6 citations
