Cross-lingual Language Model Pretraining for Retrieval
Puxuan Yu, Hongliang Fei, Ping Li
摘要
Existing research on cross-lingual retrieval cannot take good advantage of large-scale pretrained language models such as multilingual BERT and XLM. We hypothesize that the absence of cross-lingual passage-level relevance data for finetuning and the lack of querydocument style pretraining are key factors of this issue. In this paper, we introduce two novel retrieval-oriented pretraining tasks to further pretrain cross-lingual language models for downstream retrieval tasks such as cross-lingual ad-hoc retrieval (CLIR) and cross-lingual question answering (CLQA). We construct distant supervision data from multilingual Wikipedia using section alignment to support retrieval-oriented language model pretraining. We also propose to directly finetune language models on part of the evaluation collection by making Transformers capable of accepting longer sequences. Experiments on multiple benchmark datasets show that our proposed model can significantly improve upon general multilingual language models in both the cross-lingual retrieval setting and the cross-lingual transfer setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Generative Retrieval as Multi-Vector Dense RetrievalShiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen 等SIGIR 2024 · 被引用 14 次
- Soft Prompt Decoding for Multilingual Dense RetrievalZhiqi Huang, Hansi Zeng, Hamed Zamani, James AllanSIGIR 2023 · 被引用 10 次
- Improving Semantic Proximity in Information Retrieval through Cross-Lingual AlignmentSeongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper9
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 被引用 448 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 被引用 211 次
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer 等ACL 2020 · 被引用 210 次
相关 Paper
- LAReQA: Language-Agnostic Answer Retrieval from a Multilingual PoolUma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua 等EMNLP 2020 · 被引用 39 次
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentZewen Chi, Li Dong, Bo Zheng, Shaohan Huang 等ACL 2021
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 被引用 86 次
- Modeling Sequential Sentence Relation to Improve Cross-lingual Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang 等ICLR 2023 · 被引用 1 次
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 被引用 31 次
