On Bilingual Lexicon Induction with Large Language Models
Yaoyiran Li, Anna Korhonen, Ivan Vulic
Abstract
Bilingual Lexicon Induction (BLI) is a core task in multilingual NLP that still, to a large extent, relies on calculating cross-lingual word representations. Inspired by the global paradigm shift in NLP towards Large Language Models (LLMs), we examine the potential of the latest generation of LLMs for the development of bilingual lexicons. We ask the following research question: Is it possible to prompt and fine-tune multilingual LLMs (mLLMs) for BLI, and how does this approach compare against and complement current BLI approaches? To this end, we systematically study 1) zero-shot prompting for unsupervised BLI and 2) fewshot in-context prompting with a set of seed translation pairs, both without any LLM finetuning, as well as 3) standard BLI-oriented finetuning of smaller LLMs. We experiment with 18 open-source text-to-text mLLMs of different sizes (from 0.3B to 13B parameters) on two standard BLI benchmarks covering a range of typologically diverse languages. Our work is the first to demonstrate strong BLI capabilities of text-to-text mLLMs. The results reveal that few-shot prompting with in-context examples from nearest neighbours achieves the best performance, establishing new state-of-the-art BLI scores for many language pairs. We also conduct a series of in-depth analyses and ablation studies, providing more insights on BLI with (m)LLMs, also along with their limitations. Mask-Filling-Style Templates (Zero-Shot Prompting) 1 The word 'w x ' in L y is: <mask>. 2 The word w x in L y is: <mask>. 3 The word 'w x ' in L y is: <mask> 4 The word w x in L y is <mask> 5 The L x word w x in L y is: <mask>. 6 The L x word w x in L y is <mask>. 7 The L x word 'w x ' in L y is: <mask>. 8 The L x word 'w x ' in L y is <mask>. 9 The L x word w x in L y is: <mask> 10 The L x word w x in L y is <mask> 11 The L x word 'w x ' in L y is: <mask> 12 The L x word 'w x ' in L y is <mask> 13 'w x ' in L y is: <mask>. 14 w x in L y is: <mask>. 15 'w x ' in L y is: <mask> 16 w x in L y is: <mask> 17 What is the translation of the word 'w x ' into L y ? <mask>. 18 What is the translation of the word w x into L y ? <mask>. 19 What is the translation of the L x word 'w x ' into L y ? <mask>. 20 What is the translation of the L x word w x into L y ? <mask>. 21 The translation of the word 'w x ' into L y is <mask>. 22 The translation of the word w x into L y is <mask>. 23 The translation of the L x word 'w x ' into L y is <mask>. 24 How do you say 'w x ' in L y ? <mask>.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf27099c-9f81-4124-a4f1-b7ede601f030Cited by top-tier papers4
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
- DM-BLI: Dynamic Multiple Subspaces Alignment for Unsupervised Bilingual Lexicon InductionLing Hu, Yuemei XuACL 2024 · 3 citations
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao et al.ACL 2026 · 1 citation
- LexGen: Domain-aware Multilingual Lexicon GenerationAyush Maheshwari, Atul Kumar Singh, Karthika NJ, Krishnakant Bhatt et al.ACL 2025 · 1 citation
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata et al.EMNLP 2023 · 19 citations
- Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsXuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, Lidong BingACL 2024
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
- Few-shot Learning with Multilingual Generative Language ModelsXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang et al.EMNLP 2022 · 113 citations
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov et al.EMNLP 2022 · 71 citations
