LexGen: Domain-aware Multilingual Lexicon Generation
Ayush Maheshwari, Atul Kumar Singh, Karthika NJ, Krishnakant Bhatt, Preethi Jyothi, Ganesh Ramakrishnan
摘要
Lexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserving language identity. Prior work in the field primarily focuses on bilingual lexical induction, which deals with word alignments using mapping or corpora-based approaches. However, these approaches do not cater to domain-specific lexicon generation that consists of domain-specific terminology. This task becomes particularly important in specialized medical, engineering, and other technical domains, owing to the highly infrequent usage of the terms and scarcity of data involving domain-specific terms especially for low/midresource languages. In this paper, we propose a new model to generate dictionary words for 6 Indian languages in the multi-domain setting. Our model consists of domain-specific and domain-generic layers that encode information, and these layers are invoked via a learnable routing technique. We also release a new benchmark dataset consisting of >75K translation pairs across 6 Indian languages spanning 8 diverse domains. We conduct both zeroshot and few-shot experiments across multiple domains to show the efficacy of our proposed model in generalizing to unseen domains and unseen languages. Additionally, we also perform a post-hoc human evaluation on unseen languages. The source code and dataset is present at https://github.com/ Atulkmrsingh/lexgen . * Authors contributing equally. † Work done while pursuing PhD at IIT Bombay. 1 Throughout the paper, we use the terms 'lexicon' and 'dictionary' interchangeably.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 被引用 97 次
- A Graph-based Coarse-to-fine Method for Unsupervised Bilingual Lexicon InductionShuo Ren, Shujie Liu, Ming Zhou, Shuai MaACL 2020 · 被引用 13 次
- LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent SpaceTasnim Mohiuddin, M. Saiful Bari, Shafiq Rayhan JotyEMNLP 2020 · 被引用 13 次
相关 Paper
- RAPO: An Adaptive Ranking Paradigm for Bilingual Lexicon InductionZhoujin Tian, Chaozhuo Li, Shuo Ren, Zhiqiang Zuo 等EMNLP 2022 · 被引用 3 次
- XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesDhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian 等WWW 2023 · 被引用 3 次
- Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages StudyYash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi 等ACL 2021
- IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic LanguagesAman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra 等EMNLP 2022 · 被引用 16 次
- Multiple Sources are Better Than One: Incorporating External Knowledge in Low-Resource GlossingChangbing Yang, Garrett Nicolai, Miikka SilfverbergEMNLP 2024
