A Large-Scale Study of Machine Translation in Turkic Languages
Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis M. Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Behzodbek Moydinboyev, Esra Onal
摘要
Recent advances in neural machine translation (NMT) have pushed the quality of machine translation systems to the point where they are becoming widely adopted for building competitive systems. However, there is still a large number of languages that are yet to reap the benefits of NMT. In this paper, we provide the first large-scale case study of the practical application of MT in the Turkic language family in order to realize the gains of NMT for Turkic languages under high-resource to extremely low-resource scenarios. In addition to presenting an extensive analysis that identifies the bottlenecks towards building competitive systems to ameliorate data scarcity, our study has several key contributions, including, i) a large parallel corpus covering 22 Turkic languages consisting of common public datasets in combination with new datasets of approximately 2 million parallel sentences, ii) bilingual baselines for 26 language pairs, iii) novel high-quality test sets in three different translation domains and iv) human evaluation scores. All of our data, software and models are publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Glot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesAyyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini 等ACL 2023 · 被引用 14 次
- Multilingual Arbitration: Optimizing Data Pools to Accelerate Multilingual ProgressAyomide Odumakinde, Daniel D'souza, Pat Verga, Beyza Ermis 等ACL 2025 · 被引用 3 次
- Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAhmet Üstün, Viraat Aryabumi, Zheng Xin Yong, Wei-Yin Ko 等ACL 2024
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- In Neural Machine Translation, What Does Transfer Learning Transfer?Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico SennrichACL 2020 · 被引用 56 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
相关 Paper
- TUMLU: A Unified and Native Language Understanding Benchmark for Turkic LanguagesJafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova 等ACL 2025
- Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan 等ACL 2023 · 被引用 3 次
- Alternative Input Signals Ease Transfer in Multilingual Machine TranslationSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary 等ACL 2022 · 被引用 18 次
- Data and Representation for Turkish Natural Language InferenceEmrah Budur, Riza Özçelik, Tunga Gungor, Christopher PottsEMNLP 2020
- Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine TranslationTahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan 等EMNLP 2020 · 被引用 7 次
