Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AI
Octavian Alexandru Trifan, Jason Lee Weber, Marc Titus Trifan, Alexandru Nicolau, Alexander V. Veidenbaum
摘要
Edge deployment of task-oriented semantic parsers demands high accuracy under tight latency and memory budgets. We present Grammar Pruning, a lightweight zero-shot framework that begins with a user-defined schema of API calls and couples a rule-based entity extractor with an iterative grammar-constrained decoder: extracted items dynamically prune the context-free grammar, limiting generation to only those intents, slots, and values that remain plausible at each step. This aggressive searchspace reduction both reduces hallucinations and slashes decoding time. On the adapted FoodOrdering, APIMIXSNIPS, and APIMIXATIS benchmarks, Grammar Pruning with small language models achieves an average execution accuracy of over 90%-rivaling State-of-the-Art, cloud-based solutions-while sustaining at least 2x lower end-to-end latency than existing methods. By requiring nothing beyond the domain's full API schema values yet delivering precise, real-time natural-language understanding, Grammar Pruning positions itself as a practical building block for future edge-AI applications that cannot rely on large models or cloud offloading. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Grammar Prompting for Domain-Specific Language Generation with Large Language ModelsBailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao 等NeurIPS 2023 · 被引用 138 次
相关 Paper
- Rethinking Pruning for Accelerating Deep Inference At the EdgeDawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong 等KDD 2020 · 被引用 24 次
- Earley-Driven Dynamic Pruning for Efficient Structured DecodingXintong Sun, Chi Wei, Minghao Tian, Shiwen NiICML 2025
- EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation SystemsLiang Wang, Kai Wang, Ranjun Jia, Kai Lu 等ICDE 2026
- End-to-End Model Generation with Large Language Models for Adaptive IoT Application DeploymentZhenyu Wen, Jintao Feng, Nanjie Yao, Di Wu 等ICSE 2026
- EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP InferenceThierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia 等MICRO 2021 · 被引用 117 次
