Cross-lingual Retrieval for Iterative Self-Supervised Training
Chau Tran, Yuqing Tang, Xian Li, Jiatao Gu
Abstract
Recent studies have demonstrated the cross-lingual alignment ability of multilingual pretrained language models. In this work, we found that the cross-lingual alignment can be further improved by training seq2seq models on sentence pairs mined using their own encoder outputs. We utilized these findings to develop a new approachcross-lingual retrieval for iterative self-supervised training (CRISS), where mining and training processes are applied iteratively, improving cross-lingual alignment and translation ability at the same time. Using this method, we achieved stateof-the-art unsupervised machine translation results on 9 language directions with an average improvement of 2.4 BLEU, and on the Tatoeba sentence retrieval task in the XTREME benchmark on 16 languages with an average improvement of 21.5% in absolute accuracy. Furthermore, CRISS also brings an additional 1.8 BLEU improvement on average compared to mBART, when finetuned on supervised machine translation downstream tasks. Our code and pretrained models are publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c620af1-8937-4e2d-9eac-ec060fc8c00fCited by top-tier papers21
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
- ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual CorporaXuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun et al.EMNLP 2021 · 68 citations
- Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersGuanhua Chen, Shuming Ma, Yun Chen, Li Dong et al.EMNLP 2021 · 30 citations
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li et al.ICCV 2023 · 28 citations
- CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li et al.ACL 2023 · 23 citations
Builds on5
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Data-dependent Gaussian Prior Objective for Language GenerationZuchao Li, Rui Wang, Kehai Chen, Masao Utiyama et al.ICLR 2020 · 68 citations
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
Related papers
- Bilingual alignment transfers to multilingual alignment for unsupervised parallel text miningChih-chan Tien, Shane Steinert-ThrelkeldACL 2022 · 10 citations
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen et al.EMNLP 2022 · 6 citations
- Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine TranslationGuanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang et al.ACL 2022
- LAReQA: Language-Agnostic Answer Retrieval from a Multilingual PoolUma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua et al.EMNLP 2020 · 39 citations
- Cross-lingual Language Model Pretraining for RetrievalPuxuan Yu, Hongliang Fei, Ping LiWWW 2021 · 42 citations
