Grammar as Control: Modular Language Generation for the Long Tail
Ndapa Nakashole
Abstract
Large language models (LLMs) can, in principle, bootstrap language technologies for long-tail languages due to their pattern recognition capabilities. Yet in practice, without structured guidance, they produce narrow, unrepresenta-tive samples that fail to cover the morphosyn-tactic space of typologically underrepresented languages. We propose Modular Typology-Informed Generation (mTIG), a prompting framework that transforms descriptive grammars into explicit control mechanisms that guide LLMs to generate typologically balanced synthetic data for downstream training. mTIG decomposes grammars into modular grammar slices , each targeting a specific morphosyntactic phenomenon (e.g., passive voice, causative morphology). Across three low-resource languages, mTIG improves typological entropy by up to 19% and yields a “student-beats-teacher” effect, where distilled models outperform the source LLM by up to +20 chrF in machine translation. These findings show that grammar-as-control can construct training corpora wherever formal linguistic descriptions exist.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d335d61c-6e34-4987-80f7-796b6b3ed50bBuilds on1
Related papers
- Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsXuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, Lidong BingACL 2024
- GrammaMT: Improving Machine Translation with Grammar-Informed In-Context LearningRita Ramos, Everlyn Asiko Chimoto, Maartje ter Hoeve, Natalie SchluterACL 2025 · 10 citations
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeTsedeniya Kinfe Temesgen, Marion Di Marco, Alexander FraserEMNLP 2025 · 2 citations
- CLAOCS-TX: Cross-Lingual Triplet Extraction with Aspect-Opinion-Aware Code-Switched Prompting and LLM-Guided Contrastive DistillationLipika Dewangan, Chandresh Kumar MauryaACL 2026
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
