Parallel Corpus Filtering via Pre-trained Language Models
Boliang Zhang, Ajay Nagesh, Kevin Knight
Abstract
Web-crawled data provides a good source of parallel corpora for training machine translation models. It is automatically obtained, but extremely noisy, and recent work shows that neural machine translation systems are more sensitive to noise than traditional statistical machine translation methods. In this paper, we propose a novel approach to filter out noisy sentence pairs from web-crawled corpora via pre-trained language models. We measure sentence parallelism by leveraging the multilingual capability of BERT and use the Generative Pre-training (GPT) language model as a domain filter to balance data domains. We evaluate the proposed method on the WMT 2018 Parallel Corpus Filtering shared task, and on our own web-crawled Japanese-Chinese parallel corpus. Our method significantly outperforms baselines and achieves a new stateof-the-art. In an unsupervised setting, our method achieves comparable performance to the top-1 supervised method. We also evaluate on a web-crawled Japanese-Chinese parallel corpus that we make publicly available. 9 https://commoncrawl.org
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing HeuristicsAloka Fernando, Nisansa de Silva, Menan Velayuthan, Charitha Rathnayake et al.EMNLP 2025 · 1 citation
- Prevent the Language Model from being Overconfident in Neural Machine TranslationMengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou et al.ACL 2021
Related papers
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 26 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
- Breaking the Corpus Bottleneck for Context-Aware Neural Machine Translation with Cross-Task Pre-trainingLinqing Chen, Junhui Li, Zhengxian Gong, Boxing Chen et al.ACL 2021
- Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data SelectionThuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza HaffariEMNLP 2021 · 2 citations
