Dual-Alignment Pre-training for Cross-lingual Sentence Embedding
Ziheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng, Qiang Lou, Haizhen Huang, Jian Jiao, Furu Wei, Weiwei Deng, Qi Zhang
Abstract
Recent studies have shown that dual encoder models trained with the sentence-level translation ranking task are effective methods for cross-lingual sentence embedding. However, our research indicates that token-level alignment is also crucial in multilingual scenarios, which has not been fully explored previously. Based on our findings, we propose a dual-alignment pre-training (DAP) framework for cross-lingual sentence embedding that incorporates both sentence-level and token-level alignment. To achieve this, we introduce a novel representation translation learning (RTL) task, where the model learns to use one-side contextualized token representation to reconstruct its translation counterpart. This reconstruction objective encourages the model to embed translation information into the token representation. Compared to other token-level alignment methods such as translation language modeling, RTL is more suitable for dual encoder architectures and is computationally efficient. Extensive experiments on three sentencelevel cross-lingual benchmarks demonstrate that our approach can significantly improve sentence embedding. Our code is available at https://github.com/ChillingDream/DAP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f01d1dfb-8d36-4d6e-beed-add6b69686feCited by top-tier papers5
- Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMsDanni Liu, Jan NiehuesACL 2025 · 23 citations
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang et al.NeurIPS 2023 · 8 citations
- Word Alignment as Preference for Machine TranslationQiyu Wu, Masaaki Nagata, Zhongtao Miao, Yoshimasa TsuruokaEMNLP 2024 · 4 citations
- DA-Net: A Disentangled and Adaptive Network for Multi-Source Cross-Lingual Transfer LearningLing Ge, Chunming Hu, Guanghui Ma, Jihong Liu et al.AAAI 2024 · 3 citations
- MEXMA: Token-level objectives improve sentence representationsJoão Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari, Loïc BarraultACL 2025
Builds on3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Universal Sentence Representation Learning with Conditional Masked Language ModelZiyi Yang, Yinfei Yang, Daniel Cer, Jax Law et al.EMNLP 2021 · 37 citations
Related papers
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- Cross-lingual Sentence Embedding using Multi-Task LearningKoustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen et al.EMNLP 2021 · 9 citations
- Modeling Sequential Sentence Relation to Improve Cross-lingual Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ICLR 2023 · 1 citation
- Bilingual alignment transfers to multilingual alignment for unsupervised parallel text miningChih-chan Tien, Shane Steinert-ThrelkeldACL 2022 · 10 citations
- ABSent: Cross-Lingual Sentence Representation Mapping with Bidirectional GANsZuohui Fu, Yikun Xian, Shijie Geng, Yingqiang Ge et al.AAAI 2020 · 20 citations
