HintedBT: Augmenting Back-Translation with Quality and Transliteration Hints
Sahana Ramnath, Melvin Johnson, Abhirut Gupta, Aravindan Raghuveer
Abstract
Back-translation (BT) of target monolingual corpora is a widely used data augmentation strategy for neural machine translation (NMT), especially for low-resource language pairs. To improve effectiveness of the available BT data, we introduce HintedBT-a family of techniques which provides hints (through tags) to the encoder and decoder. First, we propose a novel method of using both high and low quality BT data by providing hints (as source tags on the encoder) to the model about the quality of each source-target pair. We don't filter out low quality data but instead show that these hints enable the model to learn effectively from noisy data. Second, we address the problem of predicting whether a source token needs to be translated or transliterated to the target language, which is common in crossscript translation tasks (i.e., where source and target do not share the written script). For such cases, we propose training the model with additional hints (as target tags on the decoder) that provide information about the operation required on the source (translation or both translation and transliteration). We conduct experiments and detailed analyses on standard WMT benchmarks for three crossscript low/medium-resource language pairs: Hindi,Gujarati,Tamil→English. Our methods compare favorably with five strong and well established baselines. We show that using these hints, both separately and together, significantly improves translation quality and leads to state-of-the-art performance in all three language pairs in corresponding bilingual settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3081e2c0-ae49-448d-b5a1-aa0fb2f4c131Cited by top-tier papers3
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- Back Translation for Speech-to-text Translation Without TranscriptsQingkai Fang, Yang FengACL 2023 · 9 citations
- Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time MarkersDaniel D'souza, Julia Kreutzer, Adrien Morisot, Ahmet Üstün et al.NeurIPS 2025 · 3 citations
Builds on5
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Multi-task Learning for Multilingual Neural Machine TranslationYiren Wang, ChengXiang Zhai, Hany HassanEMNLP 2020 · 58 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- Translationese as a Language in "Multilingual" NMTParker Riley, Isaac Caswell, Markus Freitag, David GrangierACL 2020
Related papers
- Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine TranslationXabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy WayACL 2020 · 1 citation
- Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning ApproachVíctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-MartínezEMNLP 2021 · 21 citations
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 75 citations
- Mirror-Generative Neural Machine TranslationZaixiang Zheng, Hao Zhou, Shujian Huang, Lei Li et al.ICLR 2020 · 37 citations
