CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn
Abstract
Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other. We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English. In addition to curating this massive dataset, we introduce baseline methods that leverage cross-lingual representations to identify aligned documents based on their textual content. Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data. Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b68cfaf-56c5-4b0d-a8e0-4bcfb07c7b26Cited by top-tier papers22
- Data Scaling Laws in NMT: The Effect of Noise and ArchitectureYamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang et al.ICML 2022 · 61 citations
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data PoisoningChang Xu, Jun Wang, Yuqing Tang, Francisco Guzmán et al.WWW 2021 · 37 citations
- Z-Code++: A Pre-trained Language Model Optimized for Abstractive SummarizationPengcheng He, Baolin Peng, Song Wang, Yang Liu et al.ACL 2023 · 27 citations
- Alternative Input Signals Ease Transfer in Multilingual Machine TranslationSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary et al.ACL 2022 · 18 citations
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled CorpusJesse Dodge, Maarten Sap, Ana Marasovic, William Agnew et al.EMNLP 2021 · 18 citations
Related papers
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
- Document-Level Machine Translation with Large-Scale Public Parallel CorporaProyag Pal, Alexandra Birch, Kenneth HeafieldACL 2024
- CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information RetrievalShuo Sun, Kevin DuhEMNLP 2020 · 53 citations
- Multimodal and Multilingual Embeddings for Large-Scale Speech MiningPaul-Ambroise Duquenne, Hongyu Gong, Holger SchwenkNeurIPS 2021 · 43 citations
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
