CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn
摘要
Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other. We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English. In addition to curating this massive dataset, we introduce baseline methods that leverage cross-lingual representations to identify aligned documents based on their textual content. Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data. Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Data Scaling Laws in NMT: The Effect of Noise and ArchitectureYamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang 等ICML 2022 · 被引用 61 次
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data PoisoningChang Xu, Jun Wang, Yuqing Tang, Francisco Guzmán 等WWW 2021 · 被引用 37 次
- Z-Code++: A Pre-trained Language Model Optimized for Abstractive SummarizationPengcheng He, Baolin Peng, Song Wang, Yang Liu 等ACL 2023 · 被引用 27 次
- Alternative Input Signals Ease Transfer in Multilingual Machine TranslationSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary 等ACL 2022 · 被引用 18 次
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled CorpusJesse Dodge, Maarten Sap, Ana Marasovic, William Agnew 等EMNLP 2021 · 被引用 18 次
相关 Paper
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave 等ACL 2021
- Document-Level Machine Translation with Large-Scale Public Parallel CorporaProyag Pal, Alexandra Birch, Kenneth HeafieldACL 2024
- CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information RetrievalShuo Sun, Kevin DuhEMNLP 2020 · 被引用 53 次
- Multimodal and Multilingual Embeddings for Large-Scale Speech MiningPaul-Ambroise Duquenne, Hongyu Gong, Holger SchwenkNeurIPS 2021 · 被引用 43 次
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
