CipherDAug: Ciphertext based Data Augmentation for Neural Machine Translation
Nishant Kambhatla, Logan Born, Anoop Sarkar
摘要
We propose a novel data-augmentation technique for neural machine translation based on ROT-k ciphertexts. ROT-k is a simple letter substitution cipher that replaces a letter in the plaintext with the kth letter after it in the alphabet. We first generate multiple ROT-k ciphertexts using different values of k for the plaintext which is the source side of the parallel data. We then leverage this enciphered training data along with the original parallel data via multi-source training to improve neural machine translation. Our method, CipherDAug, uses a co-regularization-inspired training procedure, requires no external data sources other than the original training data, and uses a standard Transformer to outperform strong data augmentation techniques on several datasets by a significant margin. This technique combines easily with existing approaches to data augmentation, and yields particularly strong results in low-resource settings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Aspect-Based Sentiment Analysis with Explicit Sentiment AugmentationsJihong Ouyang, Zhiyao Yang, Silong Liang, Bing Wang 等AAAI 2024 · 被引用 21 次
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao 等EMNLP 2022 · 被引用 20 次
- Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence ModelingZhijun Wang, Xuebo Liu, Min ZhangEMNLP 2022 · 被引用 10 次
- Curriculum Consistency Learning for Conditional Sentence GenerationLiangxin Liu, Xuebo Liu, Lian Lian, Shengjun Cheng 等EMNLP 2024 · 被引用 1 次
它引用的顶会 Paper5
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang 等NeurIPS 2021 · 被引用 610 次
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 被引用 55 次
- Sequence Generation with Mixed RepresentationsLijun Wu, Shufang Xie, Yingce Xia, Yang Fan 等ICML 2020 · 被引用 18 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
- All Word Embeddings from One EmbeddingSho Takase, Sosuke KobayashiNeurIPS 2020 · 被引用 16 次
相关 Paper
- Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning ApproachVíctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-MartínezEMNLP 2021 · 被引用 21 次
- Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng 等ACL 2022 · 被引用 26 次
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationYong Cheng, Lu Jiang, Wolfgang Macherey, Jacob EisensteinACL 2020 · 被引用 105 次
- Uncertainty-Aware Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng 等EMNLP 2020 · 被引用 20 次
