CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration
Zixin Chen, Jiachen Wang, Yumeng Li, Haobo Li, Chuhan Shi, Rong Zhang, Huamin Qu
Abstract
Grading project reports is increasingly significant in today's educational landscape, where they serve as key assessments of students' comprehensive problem-solving abilities. However, it remains challenging due to the multifaceted evaluation criteria involved, such as creativity and peer-comparative achievement. Meanwhile, instructors often struggle to maintain fairness throughout the timeconsuming grading process. Recent advances in AI, particularly large language models, have demonstrated potential for automating simpler grading tasks, such as assessing quizzes or basic writing quality. However, these tools often fall short when it comes to complex metrics, like design innovation and the practical application of knowledge, that require an instructor's educational insights and contextual understanding of the class. To address this challenge, we conducted a formative study with six instructors and developed CoGrader, which introduces a novel grading workflow combining human-LLM collaborative metrics design, benchmarking, and AI-assisted feedback. CoGrader was found effective in improving grading efficiency and consistency while providing reliable peercomparative feedback to students. We also discuss design insights and ethical considerations for the development of human-AI collaborative grading systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a940e74-ee4a-4a4c-b8a9-bbfeaa555808Builds on21
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge WorkersHao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos et al.CHI 2025 · 690 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator NeedsMajeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley et al.CHI 2024 · 246 citations
- Human Creativity in the Age of LLMs: Randomized Experiments on Divergent and Convergent ThinkingHarsh Kumar, Jonathan Vincentius, Ewan Jordan, Ashton AndersonCHI 2025 · 107 citations
Related papers
- Co-designing Large Language Model Tools for Project-Based Learning with K12 EducatorsPrerna Ravi, John Masla, Gisella Kakoti, Grace C. Lin et al.CHI 2025 · 26 citations
- Open-ended Structured Question Assessment with Human-LLM CollaborationFengyan Lin, Yanna Lin, Kai Cao, Zikun Deng et al.CHI 2026 · 1 citation
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer ReviewYaohui Zhang, Haijing Zhang, Wenlong Ji, Tianyu Hua et al.NeurIPS 2025 · 15 citations
- CogBench: a large language model walks into a psychology labJulian Coda-Forno, Marcel Binz, Jane X. Wang, Eric SchulzICML 2024 · 60 citations
- Charting the Future of AI in Project-Based Learning: A Co-Design Exploration with StudentsChengbo Zheng, Kangyu Yuan, Bingcan Guo, Reza Hadi Mogavi et al.CHI 2024 · 65 citations
