Open-ended Structured Question Assessment with Human-LLM Collaboration
Fengyan Lin, Yanna Lin, Kai Cao, Zikun Deng, Yi Cai
Abstract
Open-ended Structured Questions (OSQs) assess not only students’ knowledge but also their reasoning and expression. However, grading OSQ requires fine-grained, scoring point–level analysis, which is labor-intensive and difficult to scale. Although recent LLM-based and human–AI collaborative grading systems improve efficiency, they mainly operate at the whole-response level and lack support for point-level inspection, correction, and feedback integration. We present VeriGrader, a novel human–AI collaborative system for OSQ grading. It combines chain-of-thought prompting with scoring point– and response-level in-context learning to enable interpretable LLM grading and iterative refinement from instructor feedback. A coordinated multi-view interface supports efficient verification of response segments, matched scoring points, and rationales. We evaluate VeriGrader using real course data and a user study with 12 participants. Results show that VeriGrader improves both grading efficiency, accuracy, and consistency over the baselines, demonstrating the effectiveness of VeriGrader and promoting human–AI collaboration in educational assessment.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3dfb1f22-7a33-4fb3-b3d5-0522f4c41c2bRelated papers
- Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a BenchmarkZhiqi Yu, Xingping Liu, Haobin Mao, Mingshuo Liu et al.ICML 2026 · 1 citation
- From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended GenerationYuxin Jiang, Yufei Wang, Qiyuan Zhang, Xingshan Zeng et al.ICLR 2026 · 5 citations
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM IntegrationZixin Chen, Jiachen Wang, Yumeng Li, Haobo Li et al.UIST 2025 · 3 citations
- Think Thrice Before You Act: Progressive Thought Refinement in Large Language ModelsChengyu Du, Jinyi Han, Yizhou Ying, Aili Chen et al.ICLR 2025
