Parallel Corpus Filtering via Pre-trained Language Models
Boliang Zhang, Ajay Nagesh, Kevin Knight
摘要
Web-crawled data provides a good source of parallel corpora for training machine translation models. It is automatically obtained, but extremely noisy, and recent work shows that neural machine translation systems are more sensitive to noise than traditional statistical machine translation methods. In this paper, we propose a novel approach to filter out noisy sentence pairs from web-crawled corpora via pre-trained language models. We measure sentence parallelism by leveraging the multilingual capability of BERT and use the Generative Pre-training (GPT) language model as a domain filter to balance data domains. We evaluate the proposed method on the WMT 2018 Parallel Corpus Filtering shared task, and on our own web-crawled Japanese-Chinese parallel corpus. Our method significantly outperforms baselines and achieves a new stateof-the-art. In an unsupervised setting, our method achieves comparable performance to the top-1 supervised method. We also evaluate on a web-crawled Japanese-Chinese parallel corpus that we make publicly available. 9 https://commoncrawl.org
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing HeuristicsAloka Fernando, Nisansa de Silva, Menan Velayuthan, Charitha Rathnayake 等EMNLP 2025 · 被引用 1 次
- Prevent the Language Model from being Overconfident in Neural Machine TranslationMengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou 等ACL 2021
相关 Paper
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 被引用 26 次
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave 等ACL 2021
- Breaking the Corpus Bottleneck for Context-Aware Neural Machine Translation with Cross-Task Pre-trainingLinqing Chen, Junhui Li, Zhengxian Gong, Boxing Chen 等ACL 2021
- Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data SelectionThuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza HaffariEMNLP 2021 · 被引用 2 次
