CipherDAug: Ciphertext based Data Augmentation for Neural Machine Translation
Nishant Kambhatla, Logan Born, Anoop Sarkar
Abstract
We propose a novel data-augmentation technique for neural machine translation based on ROT-k ciphertexts. ROT-k is a simple letter substitution cipher that replaces a letter in the plaintext with the kth letter after it in the alphabet. We first generate multiple ROT-k ciphertexts using different values of k for the plaintext which is the source side of the parallel data. We then leverage this enciphered training data along with the original parallel data via multi-source training to improve neural machine translation. Our method, CipherDAug, uses a co-regularization-inspired training procedure, requires no external data sources other than the original training data, and uses a standard Transformer to outperform strong data augmentation techniques on several datasets by a significant margin. This technique combines easily with existing approaches to data augmentation, and yields particularly strong results in low-resource settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a28be9f-f481-4792-8e11-931c164a8a11Cited by top-tier papers4
- Aspect-Based Sentiment Analysis with Explicit Sentiment AugmentationsJihong Ouyang, Zhiyao Yang, Silong Liang, Bing Wang et al.AAAI 2024 · 21 citations
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao et al.EMNLP 2022 · 20 citations
- Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence ModelingZhijun Wang, Xuebo Liu, Min ZhangEMNLP 2022 · 10 citations
- Curriculum Consistency Learning for Conditional Sentence GenerationLiangxin Liu, Xuebo Liu, Lian Lian, Shengjun Cheng et al.EMNLP 2024 · 1 citation
Builds on5
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang et al.NeurIPS 2021 · 610 citations
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 55 citations
- Sequence Generation with Mixed RepresentationsLijun Wu, Shufang Xie, Yingce Xia, Yang Fan et al.ICML 2020 · 18 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
- All Word Embeddings from One EmbeddingSho Takase, Sosuke KobayashiNeurIPS 2020 · 16 citations
Related papers
- Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning ApproachVíctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-MartínezEMNLP 2021 · 21 citations
- Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng et al.ACL 2022 · 26 citations
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationYong Cheng, Lu Jiang, Wolfgang Macherey, Jacob EisensteinACL 2020 · 105 citations
- Uncertainty-Aware Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng et al.EMNLP 2020 · 20 citations
