Jump-Starting Item Parameters for Adaptive Language Tests
Arya D. McCarthy, Kevin P. Yancey, Geoffrey T. LaFlair, Jesse Egbert, Manqian Liao, Burr Settles
Abstract
A challenge in designing high-stakes language assessments is calibrating the test item difficulties, either a priori or from limited pilot test data. While prior work has addressed 'cold start' estimation of item difficulties without piloting, we devise a multi-task generalized linear model with BERT features to jump-start these estimates, rapidly improving their quality with as few as 500 test-takers and a small sample of item exposures (≈6 each) from a large item bank (≈4,000 items). Our joint model provides a principled way to compare test-taker proficiency, item difficulty, and language proficiency frameworks like the Common European Framework of Reference (CEFR). This also enables new item difficulty estimates without piloting them first, which in turn limits item exposure and thus enhances test security. Finally, using operational data from the Duolingo English Test, a high-stakes English proficiency test, we find that difficulty estimates derived using this method correlate strongly with lexico-grammatical features that correlate with reading complexity. * Research conducted during an internship at Duolingo. C1 C2 A1 A2 B1 B2 a b i l i t y d i f fi c u l t y (a) Easy for test-taker C1 C2 A1 A2 B1 B2 a b i l i t y d i f fi c u l t y (b) Hard for test-taker
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- CEFR-Based Sentence Difficulty Annotation and AssessmentYuki Arase, Satoru Uchida, Tomoyuki KajiwaraEMNLP 2022 · 17 citations
- SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty PredictionAlexander Scarlatos, Nigel Fernandez, Christopher Ormerod, Susan Lottridge et al.EMNLP 2025
- Reliable and Efficient Amortized Model-based EvaluationSang T. Truong, Yuheng Tu, Percy Liang, Bo Li et al.ICML 2025
- UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency AssessmentJoseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens et al.EMNLP 2025 · 2 citations
- Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling EstimationSang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi KoyejoICML 2026
