TeaForN: Teacher-Forcing with N-grams
Sebastian Goodman, Nan Ding, Radu Soricut
摘要
Sequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps. Our proposed method, Teacher-Forcing with N-grams (TeaForN), addresses both these problems directly, through the use of a stack of N decoders trained to decode along a secondary time axis that allows model parameter updates based on N prediction steps. TeaForN can be used with a wide class of decoder architectures and requires minimal modifications from a standard teacher-forcing setup. Empirically, we show that TeaForN boosts generation quality on one Machine Translation benchmark, WMT 2014 English-French, and two News Summarization benchmarks, CNN/Dailymail and Gigaword.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Scheduled Sampling Based on Decoding Steps for Neural Machine TranslationYijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu 等EMNLP 2021 · 被引用 9 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
- Generative Regression Based Watch Time Prediction for Short-Video RecommendationHongxu Ma, Kai Tian, Tao Zhang, Xuefeng Zhang 等WWW 2026 · 被引用 6 次
- Adaptive Bridge between Training and Inference for Dialogue GenerationHaoran Xu, Hainan Zhang, Yanyan Zou, Hongshen Chen 等EMNLP 2021 · 被引用 5 次
- ProofInfer: Generating Proof via Iterative Hierarchical InferenceZichu Fei, Qi Zhang, Xin Zhou, Tao Gui 等EMNLP 2022
它引用的顶会 Paper1
相关 Paper
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine TranslationChenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng 等AAAI 2020 · 被引用 95 次
- Multi-Granularity Optimization for Non-Autoregressive TranslationYafu Li, Leyang Cui, Yongjing Yin, Yue ZhangEMNLP 2022 · 被引用 9 次
- Guiding Teacher Forcing with Seer Forcing for Neural Machine TranslationYang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang 等ACL 2021
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei 等NeurIPS 2020 · 被引用 72 次
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 被引用 117 次
