KidLM: Advancing Language Models for Children - Early Insights and Future Directions
Mir Tafseer Nayeem, Davood Rafiei
摘要
Recent studies highlight the potential of large language models in creating educational tools for children, yet significant challenges remain in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. In this paper, we explore foundational steps toward the development of child-specific language models, emphasizing the necessity of high-quality pre-training data. We introduce a novel user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. Additionally, we propose a new training objective, Stratified Masking, which dynamically adjusts masking probabilities based on our domain-specific child language data, enabling models to prioritize vocabulary and concepts more suitable for children. Experimental evaluations demonstrate that our model excels in understanding lower gradelevel text, maintains safety by avoiding stereotypes, and captures children's unique preferences. Furthermore, we provide actionable insights for future research and development in child-specific language modeling. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SCENIC: A Location-based System to Foster Cognitive Development in Children During Car RidesLiuqing Chen, Yaxuan Song, Ke Lyu, Shuhong Xiao 等UIST 2025 · 被引用 2 次
- Biased Tales: Cultural and Topic Bias in Generating Children's StoriesDonya Rooein, Vilém Zouhar, Debora Nozza, Dirk HovyEMNLP 2025
- EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMsNumaan Naeem, Abdellah El Mekki, Muhammad Abdul-MageedEMNLP 2025
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos 等ICLR 2024 · 被引用 433 次
相关 Paper
- On the Automatic Generation and Simplification of Children's StoriesMaria R. Valentini, Jennifer Weber, Jesus Salcido, Téa Wright 等EMNLP 2023 · 被引用 7 次
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza 等ICLR 2026 · 被引用 22 次
- Safety Pretraining: Toward the Next Generation of Safe AIPratyush Maini, Sachin Goyal, Dylan Sam, Alexander Robey 等NeurIPS 2025 · 被引用 50 次
- Successor Features for Efficient Multi-Subject Controlled Text GenerationMeng Cao, Mehdi Fatemi, Jackie C. K. Cheung, Samira ShabanianICML 2024 · 被引用 2 次
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan 等EMNLP 2021 · 被引用 129 次
