KidLM: Advancing Language Models for Children - Early Insights and Future Directions
Mir Tafseer Nayeem, Davood Rafiei
Abstract
Recent studies highlight the potential of large language models in creating educational tools for children, yet significant challenges remain in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. In this paper, we explore foundational steps toward the development of child-specific language models, emphasizing the necessity of high-quality pre-training data. We introduce a novel user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. Additionally, we propose a new training objective, Stratified Masking, which dynamically adjusts masking probabilities based on our domain-specific child language data, enabling models to prioritize vocabulary and concepts more suitable for children. Experimental evaluations demonstrate that our model excels in understanding lower gradelevel text, maintains safety by avoiding stereotypes, and captures children's unique preferences. Furthermore, we provide actionable insights for future research and development in child-specific language modeling. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d25d883-a7af-48e6-9f5d-c08ee52b4345Cited by top-tier papers3
- SCENIC: A Location-based System to Foster Cognitive Development in Children During Car RidesLiuqing Chen, Yaxuan Song, Ke Lyu, Shuhong Xiao et al.UIST 2025 · 2 citations
- Biased Tales: Cultural and Topic Bias in Generating Children's StoriesDonya Rooein, Vilém Zouhar, Debora Nozza, Dirk HovyEMNLP 2025
- EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMsNumaan Naeem, Abdellah El Mekki, Muhammad Abdul-MageedEMNLP 2025
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos et al.ICLR 2024 · 433 citations
Related papers
- On the Automatic Generation and Simplification of Children's StoriesMaria R. Valentini, Jennifer Weber, Jesus Salcido, Téa Wright et al.EMNLP 2023 · 7 citations
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza et al.ICLR 2026 · 22 citations
- Safety Pretraining: Toward the Next Generation of Safe AIPratyush Maini, Sachin Goyal, Dylan Sam, Alexander Robey et al.NeurIPS 2025 · 50 citations
- Successor Features for Efficient Multi-Subject Controlled Text GenerationMeng Cao, Mehdi Fatemi, Jackie C. K. Cheung, Samira ShabanianICML 2024 · 2 citations
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan et al.EMNLP 2021 · 129 citations
