Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
Nurkhan Laiyk, Daniil Orel, Rituraj Joshi, Maiya Goloburda, Yuxia Wang, Preslav Nakov, Fajri Koto
摘要
Instruction tuning in low-resource languages remains underexplored due to limited text data, particularly in government and cultural domains. To address this, we introduce and open-source a large-scale (10,600 samples) instruction-following (IFT) dataset, 1 covering key institutional and cultural knowledge relevant to Kazakhstan. Our dataset enhances LLMs' understanding of procedural, legal, and structural governance topics. We employ LLM-assisted data generation, comparing openweight and closed-weight models for dataset construction, and select GPT-4o as the backbone. Each entity of our dataset undergoes full manual verification to ensure high quality. We also show that fine-tuning Qwen, Falcon, and Gemma on our dataset leads to consistent performance improvements in both multiplechoice and generative tasks, demonstrating the potential of LLM-assisted instruction tuning for low-resource languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Unnatural Instructions: Tuning Language Models with (Almost) No Human LaborOr Honovich, Thomas Scialom, Omer Levy, Timo SchickACL 2023 · 被引用 92 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 等ACL 2024 · 被引用 30 次
相关 Paper
- Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language ModelsGuanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia 等ICLR 2025 · 被引用 4 次
- SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-ReflectionLiangxin Liu, Xuebo Liu, Derek F. Wong, Dongfang Li 等NeurIPS 2024 · 被引用 49 次
- Scaling Towards the Information Boundary of Instructions through Data SynthesizingLi Du, Hanyu Zhao, Yiming Ju, Tengfei PanAAAI 2026
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of KazakhstanMukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov 等ACL 2025 · 被引用 10 次
