Scheduled Sampling Based on Decoding Steps for Neural Machine Translation
Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, Jie Zhou
摘要
Scheduled sampling is widely used to mitigate the exposure bias problem for neural machine translation. Its core motivation is to simulate the inference scene during training by replacing ground-truth tokens with predicted tokens, thus bridging the gap between training and inference. However, vanilla scheduled sampling is merely based on training steps and equally treats all decoding steps. Namely, it simulates an inference scene with uniform error rates, which disobeys the real inference scene, where larger decoding steps usually have higher error rates due to error accumulations. To alleviate the above discrepancy, we propose scheduled sampling methods based on decoding steps, increasing the selection chance of predicted tokens with the growth of decoding steps. Consequently, we can more realistically simulate the inference scene during training, thus better bridging the gap between training and inference. Moreover, we investigate scheduled sampling based on both training steps and decoding steps for further improvements. Experimentally, our approaches significantly outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks. Additionally, our approaches also generalize well to the text summarization task on two popular benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen 等ICCV 2023 · 被引用 35 次
- Reducing Position Bias in Simultaneous Machine Translation with Length-Aware FrameworkShaolei Zhang, Yang FengACL 2022 · 被引用 23 次
- Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine TranslationSongming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen 等ACL 2022 · 被引用 13 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
- A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingNitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir KantorACL 2023 · 被引用 5 次
它引用的顶会 Paper4
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 被引用 88 次
- Better Fine-Tuning by Reducing Representational CollapseArmen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal 等ICLR 2021 · 被引用 20 次
- TeaForN: Teacher-Forcing with N-gramsSebastian Goodman, Nan Ding, Radu SoricutEMNLP 2020 · 被引用 1 次
相关 Paper
- Mixed Cross Entropy Loss for Neural Machine TranslationHaoran Li, Wei LuICML 2021 · 被引用 21 次
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu 等EMNLP 2020 · 被引用 11 次
- Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion ModelsZhiyao Ren, Yibing Zhan, Liang Ding, Gaoang Wang 等AAAI 2024 · 被引用 15 次
- Understanding and Bridging the Modality Gap for Speech TranslationQingkai Fang, Yang FengACL 2023 · 被引用 12 次
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 被引用 117 次
