LexGen: Domain-aware Multilingual Lexicon Generation
Ayush Maheshwari, Atul Kumar Singh, Karthika NJ, Krishnakant Bhatt, Preethi Jyothi, Ganesh Ramakrishnan
Abstract
Lexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserving language identity. Prior work in the field primarily focuses on bilingual lexical induction, which deals with word alignments using mapping or corpora-based approaches. However, these approaches do not cater to domain-specific lexicon generation that consists of domain-specific terminology. This task becomes particularly important in specialized medical, engineering, and other technical domains, owing to the highly infrequent usage of the terms and scarcity of data involving domain-specific terms especially for low/midresource languages. In this paper, we propose a new model to generate dictionary words for 6 Indian languages in the multi-domain setting. Our model consists of domain-specific and domain-generic layers that encode information, and these layers are invoked via a learnable routing technique. We also release a new benchmark dataset consisting of >75K translation pairs across 6 Indian languages spanning 8 diverse domains. We conduct both zeroshot and few-shot experiments across multiple domains to show the efficacy of our proposed model in generalizing to unseen domains and unseen languages. Additionally, we also perform a post-hoc human evaluation on unseen languages. The source code and dataset is present at https://github.com/ Atulkmrsingh/lexgen . * Authors contributing equally. † Work done while pursuing PhD at IIT Bombay. 1 Throughout the paper, we use the terms 'lexicon' and 'dictionary' interchangeably.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbd4e015-cc9e-4352-91d6-a8503c4374d9Builds on9
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 97 citations
- A Graph-based Coarse-to-fine Method for Unsupervised Bilingual Lexicon InductionShuo Ren, Shujie Liu, Ming Zhou, Shuai MaACL 2020 · 13 citations
- LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent SpaceTasnim Mohiuddin, M. Saiful Bari, Shafiq Rayhan JotyEMNLP 2020 · 13 citations
Related papers
- RAPO: An Adaptive Ranking Paradigm for Bilingual Lexicon InductionZhoujin Tian, Chaozhuo Li, Shuo Ren, Zhiqiang Zuo et al.EMNLP 2022 · 3 citations
- XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesDhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian et al.WWW 2023 · 3 citations
- Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages StudyYash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi et al.ACL 2021
- IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic LanguagesAman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra et al.EMNLP 2022 · 16 citations
- Multiple Sources are Better Than One: Incorporating External Knowledge in Low-Resource GlossingChangbing Yang, Garrett Nicolai, Miikka SilfverbergEMNLP 2024
