Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs
Hantao Hua, Jiming Su, hao tang, Yiping Yao, Feng Zhu
摘要
Grammar-constrained decoding enables large language models (LLMs) to reliably generate structured outputs such as JSON, SQL, and domain-specific programs. Existing systems often enforce constraints by executing byte-level parser logic inside the token-level decoding loop, introducing CPU-side control flow and CPU--GPU synchronization that become bottlenecks under continuous batching. We propose Gram2Token, a GPU-native framework that preprocesses deterministic byte-level grammar execution into token-level transitions before inference. Gram2Token aligns tokenizer byte sequences with grammar transitions through a trie and groups tokens with identical transition outcomes across preprocessed grammar states. These categories yield compact validity masks and transition tables, reducing run-time enforcement to category lookup, masking, and state update rather than parser-style byte traversal. Across four model families under schema-diverse continuous batching, Gram2Token achieves a geometric-mean throughput improvement of 1.38 over the strongest baseline, with a maximum speedup of 1.85, at the cost of additional preprocessing and time-to-first-token overhead. Break-even and grammar-complexity analyses show that this overhead is amortized by grammar reuse, longer outputs, and larger batches. These results show that token-level grammar preprocessing is an effective design point for high-throughput structured LLM serving. Code is available at https://github.com/Paradozile/Gram2Token.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Large Language Models as Tool MakersTianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen 等ICLR 2024 · 被引用 283 次
- Grammar-Aligned DecodingKanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova 等NeurIPS 2024 · 被引用 73 次
相关 Paper
- Efficient Grammar-Constrained Decoding via Parser Stack ClassificationYongmin Li, Yihong Dong, Jia Li, Ge LiISSTA 2026
- Flexible and Efficient Grammar-Constrained DecodingKanghee Park, Timothy Zhou, Loris D'AntoniICML 2025
- Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM GenerationJunyi Chen, Shihao Bai, Zaijun Wang, Siyu Wu 等ACL 2025
- Earley-Driven Dynamic Pruning for Efficient Structured DecodingXintong Sun, Chi Wei, Minghao Tian, Shiwen NiICML 2025
- The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained DecodingAvinash Reddy, Thayne Walker, Jaime Ide, Amrit Singh BediICML 2026 · 被引用 6 次
