Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding
Xin Sun, Tao Ge, Furu Wei, Houfeng Wang
Abstract
In this paper, we propose Shallow Aggressive Decoding (SAD) to improve the online inference efficiency of the Transformer for instantaneous Grammatical Error Correction (GEC). SAD optimizes the online inference efficiency for GEC by two innovations: 1) it aggressively decodes as many tokens as possible in parallel instead of always decoding only one token in each step to improve computational parallelism; 2) it uses a shallow decoder instead of the conventional Transformer architecture with balanced encoder-decoder depth to reduce the computational cost during inference. Experiments in both English and Chinese GEC benchmarks show that aggressive decoding could yield the same predictions as greedy decoding but with a significant speedup for online inference. Its combination with the shallow decoder could offer an even higher online inference speedup over the powerful Transformer baseline without quality loss. Not only does our approach allow a single model to achieve the state-of-the-art results in English GEC benchmarks: 66.4 F 0.5 in the CoNLL-14 and 72.9 F 0.5 in the BEA-19 test set with an almost 10× online inference speedup over the Transformer-big model, but also it is easily adapted to other languages. Our code is available at https://github.com/AutoTemp/ Shallow-Aggressive-Decoding .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67c5da75-c578-4b18-9c2c-c12e90e43829Cited by top-tier papers26
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 347 citations
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 143 citations
- UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data ScienceYazheng Yang, Yuqi Wang, Guang Liu, Ledell Wu et al.ICLR 2024 · 35 citations
Builds on2
Related papers
- GEC-DePenD: Non-Autoregressive Grammatical Error Correction with Decoupled Permutation and DecodingKonstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko et al.ACL 2023 · 4 citations
- MaskGEC: Improving Neural Grammatical Error Correction via Dynamic MaskingZewei Zhao, Houfeng WangAAAI 2020 · 71 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Glancing Transformer for Non-Autoregressive Neural Machine TranslationLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang et al.ACL 2021
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong et al.ACL 2023 · 21 citations
