Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics
Aloka Fernando, Nisansa de Silva, Menan Velayuthan, Charitha Rathnayake, Surangika Ranathunga
Abstract
Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from webmined corpora. Ranking sentence pairs using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) is the most common PDC technique. However, previous research has shown that the choice of the multiPLM significantly impacts the quality of the filtered parallel corpus, and the Neural Machine Translation (NMT) models trained using such data show a disparity across multiPLMs. This paper shows that this disparity is due to different multiPLMs being biased towards certain types of sentence pairs, which are treated as noise from an NMT point of view. We show that such noisy parallel sentences can be removed to a certain extent by employing a series of heuristics. The NMT models, trained using the curated corpus, lead to producing better results while minimizing the disparity across multiPLMs. We publicly release the source code and the curated datasets 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80da3696-c851-4e0d-a4fc-61c14aa97002Builds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Parallel Corpus Filtering via Pre-trained Language ModelsBoliang Zhang, Ajay Nagesh, Kevin KnightACL 2020 · 17 citations
- CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsAhmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp KoehnEMNLP 2020 · 6 citations
Related papers
- Analyzing Similarity Metrics for Data Selection for Language Model PretrainingDylan Sam, Ayan Chakrabarti, Afshin Rostamizadeh, Srikumar Ramalingam et al.NeurIPS 2025 · 3 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 26 citations
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationNattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto OnizukaEMNLP 2021 · 16 citations
