WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
Peerat Limkonchotiwat, Pume Tuchinda, Lalita Lowphansirikul, Surapon Nonesung, Panuthep Tasawong, Alham Fikri Aji, Can Udomcharoenchaikit, Sarana Nutanong
Abstract
Large language models excel at instructionfollowing in English, but their performance in low-resource languages like Thai remains underexplored. Existing benchmarks often rely on translations, missing cultural and domainspecific nuances needed for real-world use. We present ThaiInstruct, a human-authored Thai dataset for evaluation and instruction tuning, covering four professional domains and seven task types. Created through a multi-stage quality control process with annotators, domain experts, and AI researchers, ThaiInstruct supports two studies: (1) a zero-shot evaluation showing performance gaps on culturally and professionally specific tasks, and (2) an instruction tuning study with ablations isolating the effect of native supervision. Models fine-tuned on Thai-Instruct outperform those using translated data in both in-domain and out-of-domain benchmarks. These findings underscore the need for culturally and professionally grounded instruction data to improve LLM alignment in lowresource, linguistically diverse settings. 12
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 362c28cd-0daf-4fab-89b4-4aa847fa9984Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian LanguagesHoly Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James V. Miranda et al.EMNLP 2024 · 8 citations
- Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?Nishant Balepur, Abhilasha Ravichander, Rachel RudingerACL 2024 · 5 citations
Related papers
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng et al.ACL 2025 · 3 citations
- Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?Pinzhen Chen, Simon Yu, Zhicheng Guo, Barry HaddowEMNLP 2024 · 4 citations
- LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language ModelsJian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao et al.ACL 2026 · 1 citation
- DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-TuningAnum Afzal, Florian Matthes, Alexander R. FabbriEMNLP 2025
- InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningPrakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri et al.EMNLP 2022 · 26 citations
