Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency
Eric Zelikman, Wanjing Anya Ma, Jasmine E. Tran, Diyi Yang, Jason D. Yeatman, Nick Haber
Abstract
Developing an educational test can be expensive and time-consuming, as each item must be written by experts and then evaluated by collecting hundreds of student responses. Moreover, many tests require multiple distinct sets of questions administered throughout the school year to closely monitor students’ progress, known as parallel tests. In this study, we focus on tests of silent sentence reading efficiency, used to assess students’ reading ability over time. To generate high-quality parallel tests, we propose to fine-tune large language models (LLMs) to simulate how previous students would have responded to unseen items. With these simulated responses, we can estimate each item’s difficulty and ambiguity. We first use GPT-4 to generate new test items following a list of expert-developed rules and then apply a fine-tuned LLM to filter the items based on criteria from psychological measurements. We also propose an optimal-transport-inspired technique for generating parallel tests and show the generated tests closely correspond to the original test’s difficulty and reliability based on crowdworker responses. Our evaluation of a generated test with 234 students from grades 2 to 8 produces test scores highly correlated (r=0.93) to those of a standard test form written by human experts and evaluated across thousands of K-12 students.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66302f3f-3ef0-40a4-a4ed-fcd9c4893687Cited by top-tier papers4
- RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward RedistributionJiahui Li, Lin Li, Tai-Wei Chang, Kun Kuang et al.EMNLP 2025
- SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty PredictionAlexander Scarlatos, Nigel Fernandez, Christopher Ormerod, Susan Lottridge et al.EMNLP 2025
- Uncertainty Quantification for LLM-Based Survey SimulationsChengpiao Huang, Yuhang Wu, Kaizheng WangICML 2025
- SOCIAL SCAFFOLDS: A Generalization Framework for Social Understanding TasksRitam Dutt, Carolyn P. Rosé, Maarten SapEMNLP 2025
Builds on1
Related papers
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit TokensYinhan He, Wendy Zheng, Yaochen Zhu, Zaiyi Zheng et al.NeurIPS 2025 · 19 citations
- Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?Zekai Shao, Siyu Yuan, Lin Gao, Yixuan He et al.CHI 2025 · 12 citations
- Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student CourseCheng-Han Chiang, Wei-Chih Chen, Chun-Yi Kuan, Chienchou Yang et al.EMNLP 2024 · 14 citations
- Finetuning LLMs for Human Behavior Prediction in Social Science ExperimentsAkaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. BernsteinEMNLP 2025 · 12 citations
