A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
Tatiana Anikina, Ján Cegin, Jakub Simko, Simon Ostermann
Abstract
Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models. However, a comparison of various generation strategies for low-resource language settings is lacking. While various prompting strategies have been proposed-such as demonstrations, labelbased summaries, and self-revision-their comparative effectiveness remains unclear, especially for low-resource languages. In this paper, we systematically evaluate the performance of these generation strategies and their combinations across 11 typologically diverse languages, including several extremely low-resource ones. Using three NLP tasks and four open-source LLMs, we assess downstream model performance on generated versus gold-standard data. Our results show that strategic combinations of generation methods-particularly targetlanguage demonstrations with LLM-based revisions-yield strong performance, narrowing the gap with real data to as little as 5% in some settings. We also find that smart prompting techniques can reduce the advantage of larger LLMs, highlighting efficient generation strategies for synthetic data generation in lowresource scenarios with smaller models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d592bea-1220-4278-8562-d66601ee7845Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
- ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model RobustnessJán Cegin, Jakub Simko, Peter BrusilovskyEMNLP 2023 · 25 citations
- DALE: Generative Data Augmentation for Low-Resource Legal NLPSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. et al.EMNLP 2023 · 10 citations
- People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language DetectionIndira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein et al.EMNLP 2023 · 7 citations
Related papers
- Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsXuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, Lidong BingACL 2024
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo et al.EMNLP 2025 · 2 citations
- UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic LanguagesPranjal A. Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali et al.ACL 2026 · 1 citation
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
- Pretraining Language Models Using TranslationeseMeet Doshi, Raj Dabre, Pushpak BhattacharyyaEMNLP 2024
