Open-ended Structured Question Assessment with Human-LLM Collaboration
Fengyan Lin, Yanna Lin, Kai Cao, Zikun Deng, Yi Cai
摘要
Open-ended Structured Questions (OSQs) assess not only students’ knowledge but also their reasoning and expression. However, grading OSQ requires fine-grained, scoring point–level analysis, which is labor-intensive and difficult to scale. Although recent LLM-based and human–AI collaborative grading systems improve efficiency, they mainly operate at the whole-response level and lack support for point-level inspection, correction, and feedback integration. We present VeriGrader, a novel human–AI collaborative system for OSQ grading. It combines chain-of-thought prompting with scoring point– and response-level in-context learning to enable interpretable LLM grading and iterative refinement from instructor feedback. A coordinated multi-view interface supports efficient verification of response segments, matched scoring points, and rationales. We evaluate VeriGrader using real course data and a user study with 12 participants. Results show that VeriGrader improves both grading efficiency, accuracy, and consistency over the baselines, demonstrating the effectiveness of VeriGrader and promoting human–AI collaboration in educational assessment.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a BenchmarkZhiqi Yu, Xingping Liu, Haobin Mao, Mingshuo Liu 等ICML 2026 · 被引用 1 次
- From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended GenerationYuxin Jiang, Yufei Wang, Qiyuan Zhang, Xingshan Zeng 等ICLR 2026 · 被引用 5 次
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran 等UIST 2024 · 被引用 143 次
- CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM IntegrationZixin Chen, Jiachen Wang, Yumeng Li, Haobo Li 等UIST 2025 · 被引用 3 次
- Think Thrice Before You Act: Progressive Thought Refinement in Large Language ModelsChengyu Du, Jinyi Han, Yizhou Ying, Aili Chen 等ICLR 2025
